A data set construction method, model training method and corresponding device

By using OCR recognition technology to obtain element information and position information in the bill image, determine categories and mark them, the problems of low data collection accuracy and cumbersome annotation process in the prior art are solved, and higher bill recognition accuracy and data set accuracy are achieved.

CN114067343BActive Publication Date: 2025-05-16CHINA CONSTRUCTION BANK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111421423.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-26
Publication Date
2025-05-16
Estimated Expiration
2041-11-26

AI Technical Summary

Technical Problem

In the prior art, in the identification of bills, data collection accuracy is low, which affects the recognition accuracy. In the data sets in natural language processing, they usually lose the location information of text on the bills, resulting in a cumbersome labeling process.

Method used

By obtaining the to-processed ticket image, OCR recognition is performed to determine the information and location of the feature entity, the category of the feature entity is determined based on this information, and the corresponding label is applied to the annotation, and the data set containing the location information is constructed.

Benefits of technology

It improves the accuracy of the bill data set and the accuracy of the bill recognition model, simplifies the data annotation process, and improves the accuracy of bill recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114067343B_ABST
    Figure CN114067343B_ABST
Patent Text Reader

Abstract

The embodiment of the present application relates to the field of data processing, and in particular discloses a method for constructing a data set, a model training method, a device, an electronic device, and a storage medium. The method includes: obtaining a bill image to be processed; for each bill image, performing OCR recognition on the bill image, determining the element information of each element entity in the bill image and the location information of each element entity; the element information includes at least one of text information, table information, and signature information; for each element entity, determining the category of the element entity according to the element information and location information; based on each element entity, applying a label corresponding to the category of the element entity to annotate the element entity; determining the element set composed of each annotated element entity in each bill image to be processed as a data set. It is used to improve the accuracy of the data set in the collected bill image, and then apply the data set to bill recognition to improve the accuracy of bill recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data set construction method, a model training method, a corresponding device, an electronic device and a storage medium. Background Art

[0002] Bills are an important text carrier of structured information. With the development of social situation, the styles of bills have shown linear growth and developed into various types. When relevant departments make reimbursements, they need to review several or even dozens of different types of bills, and some bill structures have great similarities.

[0003] In the prior art, a large amount of bill data is usually collected and processed accordingly to identify the bill. Therefore, the accuracy of bill data collection directly affects the recognition accuracy. Summary of the invention

[0004] The embodiments of the present application provide a data set construction method, a model training method, a corresponding device, an electronic device and a storage medium, which are used to improve the accuracy of the data set in the collected bill images, and then apply the data set to bill recognition to improve the accuracy of bill recognition.

[0005] In a first aspect, an embodiment of the present application provides a method for constructing a data set, including:

[0006] Get the bill image to be processed;

[0007] For each of the bill images, perform OCR recognition on the bill image to determine the element information of each element entity in the bill image and the position information of each element entity; wherein the element information includes at least one of text information, table information and signature information;

[0008] For each element entity, determining a category of the element entity according to the element information and the location information;

[0009] Based on each feature entity, label the feature entity using a label corresponding to the category of the feature entity;

[0010] A set of elements formed by each annotated element entity in each of the bill images to be processed is determined as a data set.

[0011] In some exemplary implementations, determining the category of the feature entity according to the feature information and the location information includes:

[0012] If the element information does not include preset keyword information, the adjacent element entity of the element entity is determined according to the location information of the element information, and the category of the element entity is determined according to the adjacent element entity; wherein the distance between the adjacent element entity and the element entity on the bill image is less than a preset distance threshold.

[0013] In some exemplary embodiments, based on the category of each feature entity, labeling the feature entity using a label corresponding to the category includes:

[0014] If the category of the feature entity is a simple category, the feature entity is marked as a label of the feature entity; wherein the number of feature entities in the area where the feature entity of the simple category is located is one;

[0015] If the category of the feature entity is a composite category, the feature entity is marked as a composite label; wherein the composite label includes labels of each feature entity in the composite feature and corresponding feature values; and the number of feature entities in the area where the feature entity of the composite category is located is at least two.

[0016] In some exemplary embodiments, before labeling the feature entity based on the category of each feature entity and applying the label corresponding to the category, the method further includes:

[0017] Display each element entity according to the preset display form.

[0018] In some exemplary implementations, if the preset display format is HTML format, then displaying each element entity according to the preset display format includes:

[0019] Determine an element text box, and display the element text box in HTML format; wherein the element text box includes attribute information of the element to be displayed;

[0020] If the preset display format is in JSON format, then displaying each element entity according to the preset display format includes:

[0021] The element key-value pairs are displayed according to the determined correspondence between the element key-value pairs; wherein the key in the key-value pair is the serial number of the element to be displayed, and the value in the key-value pair is the nested JSON string of the element to be displayed.

[0022] In some exemplary embodiments, before determining the element text box and displaying the element text box in HTML format, the method further includes:

[0023] If the element to be displayed is an element in a table, then determining the attribute label of the element text box;

[0024] And the attribute label is displayed as the auxiliary display information of the element text box.

[0025] In a second aspect, an embodiment of the present application provides a method for training a bill recognition model, comprising:

[0026] Acquire a training data set, wherein the training data includes a data set obtained by applying the method described in the first aspect above;

[0027] The pre-built neural network model is trained using the data set until the neural network model converges to obtain a bill recognition model.

[0028] In a third aspect, an embodiment of the present application provides a device for constructing a data set, including:

[0029] An image acquisition module, used to acquire the image of the bill to be processed;

[0030] An image recognition module, for performing OCR recognition on each of the bill images, and determining element information of each element entity in the bill image and position information of each element entity; wherein the element information includes at least one of text information, table information and signature information;

[0031] A category determination module, used for determining the category of each element entity according to the element information and the location information;

[0032] A labeling module, used for labeling each feature entity by applying a label corresponding to the category of the feature entity;

[0033] The data set determination module is used to determine that the element set composed of each annotated element entity in each of the bill images to be processed is a data set.

[0034] In some exemplary embodiments, the category determination module is specifically configured to:

[0035] If the element information does not include preset keyword information, the adjacent element entity of the element entity is determined according to the location information of the element information, and the category of the element entity is determined according to the adjacent element entity; wherein the distance between the adjacent element entity and the element entity on the bill image is less than a preset distance threshold.

[0036] In some exemplary embodiments, the annotation module is specifically used to:

[0037] If the category of the feature entity is a simple category, the feature entity is marked as a label of the feature entity; wherein the number of feature entities in the area where the feature entity of the simple category is located is one;

[0038] If the category of the feature entity is a composite category, the feature entity is marked as a composite label; wherein the composite label includes labels of each feature entity in the composite feature and corresponding feature values; and the number of feature entities in the area where the feature entity of the composite category is located is at least two.

[0039] In some exemplary embodiments, a display module is further included, and the display module is used for, before labeling the feature entity based on the category of each feature entity and applying the label corresponding to the category:

[0040] Display each element entity according to the preset display form.

[0041] In some exemplary embodiments, if the preset display format is HTML format, the display module is specifically used to:

[0042] Determine an element text box, and display the element text box in HTML format; wherein the element text box includes attribute information of the element to be displayed;

[0043] If the preset display format is json format, the display module is specifically used to:

[0044] The element key-value pairs are displayed according to the determined correspondence between the element key-value pairs; wherein the key in the key-value pair is the serial number of the element to be displayed, and the value in the key-value pair is the nested JSON string of the element to be displayed.

[0045] In some exemplary embodiments, the display module further includes, before determining the element text box and displaying the element text box in HTML format:

[0046] If the element to be displayed is an element in a table, then determining the attribute label of the element text box;

[0047] And the attribute label is displayed as the auxiliary display information of the element text box.

[0048] In a fourth aspect, an embodiment of the present application provides a training device for a bill recognition model, comprising:

[0049] A data set acquisition module, used to acquire a training data set, wherein the training data includes a data set obtained by the method described in the second aspect above;

[0050] The training module is used to train the pre-built neural network model using the data set until the neural network model converges to obtain a bill recognition model.

[0051] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the methods in the first or second aspects above when executing the computer program.

[0052] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the steps of any one of the methods of the first aspect or the second aspect described above.

[0053] In the seventh aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium, and at least one processor of the device reads and executes the computer program from the computer-readable storage medium, so that the device performs the method shown in any one of the embodiments of the first aspect or the second aspect.

[0054] In the embodiment of the present application, after obtaining the bill image to be processed, OCR recognition is performed on each bill image to determine the element information of each element entity in the bill image and the position information of each element entity; in this way, the category of the corresponding element entity can be determined together with the element information and the position information, and then in the process of data annotation, the element entity can be annotated with the label corresponding to the category of the element entity, and the element set composed of the annotated element entities is a data set. The data set includes not only the element information, but also the position information, thus improving the accuracy of the bill recognition model obtained by using the data set as a sample for training, and improving the accuracy of bill recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. Obviously, the drawings introduced below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0056] Figure 1 A schematic diagram of an application scenario of bill recognition provided by an embodiment of the present application;

[0057] Figure 2 A schematic diagram of a bill style provided in one embodiment of the present application;

[0058] Figure 3A flowchart of a method for constructing a data set provided in one embodiment of the present application;

[0059] Figure 4 A display effect diagram after identifying an element entity provided by an embodiment of the present application;

[0060] Figure 5 A schematic diagram of a marking process provided in one embodiment of the present application;

[0061] Figure 6 A schematic diagram marked in Excel provided in one embodiment of the present application;

[0062] Figure 7 A schematic diagram of an integrated labeling result provided in an embodiment of the present application;

[0063] Figure 8 A flowchart of a method for training a bill recognition model provided in one embodiment of the present application;

[0064] Fig. 9 A schematic diagram of the structure of a data set construction device provided in one embodiment of the present application;

[0065] Fig.10 A schematic diagram of the structure of a training device for a bill recognition model provided in one embodiment of the present application;

[0066] Fig.11 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application;

[0067] Fig.12 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0068] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0069] For ease of understanding, the terms involved in the embodiments of the present application are explained below:

[0070] (1) OCR (Optical Character Recognition): Analyze and process image files to obtain text and layout information on the image.

[0071] (2) NLP (Natural Language Processing): studies various theories and methods that can enable effective communication between humans and computers using natural language.

[0072] (3) Sequence labeling: Sequence labeling is a relatively simple NLP task that is used to solve a series of character classification problems, such as word segmentation, part-of-speech tagging, named entity recognition, and relationship extraction.

[0073] Any number of elements in the drawings is for illustrative purposes only and not limiting, and any naming is for distinction only and does not have any limiting meaning.

[0074] In the specific practice, bills are an important text carrier of structured information. With the development of the social situation, the styles of bills have shown a linear growth and developed into various types. When the relevant departments make reimbursements, they need to review several or even dozens of different types of bills, and some bill structures have great similarities. In the prior art, a large amount of bill corpus is usually collected and then processed by some methods to identify bills. Therefore, the accuracy of bill data collection directly affects the recognition accuracy.

[0075] In addition, natural language datasets are usually generated by sequence annotation, which will lose the location information of the text itself on the bill. For image text annotation, the text information on the image needs to be manually input and entity annotation is performed, which is time-consuming and laborious.

[0076] In terms of specific application scenarios, the direction of natural language processing is still mainly question-answering robots. The data set is more about semantic analysis of the sentence itself, analyzing the part of speech of the words in the sentence and the semantic relationship in the context, so as to give corresponding answers. Such annotated data is completely sufficient for question-answering robots with the goal of dialogue, but for bill information extraction where the position of the text itself is also an extraction factor, simple sequence annotation cannot well reflect the position of the entity in the bill itself, and additional annotation is required, which is more cumbersome.

[0077] To this end, the present application provides a method for constructing a data set, in which not only the element information of the element entity is used, but also the location information of each element entity is used, so that the category of the element entity can be determined according to the element information and location information, and then the element entity is labeled according to the label corresponding to the category of the element entity, and the element set composed of each labeled element entity is a data set. The accuracy of the bill recognition model obtained by using this data set as a sample training is high, and the accuracy of subsequent bill recognition is also high.

[0078] After introducing the design ideas of the embodiments of the present application, the following briefly introduces the application scenarios to which the technical solutions of the embodiments of the present application can be applied. It should be noted that the application scenarios introduced below are only used to illustrate the embodiments of the present application and are not limited. In specific implementation, the technical solutions provided by the embodiments of the present application can be flexibly applied according to actual needs.

[0079] refer to Figure 1 , which is a schematic diagram of an application scenario of bill recognition provided by an embodiment of the present application. The application scenario includes multiple terminal devices 101 (including terminal device 101-1, terminal device 101-2, ... terminal device 101-n) and a server 102. Among them, the terminal device 101 and the server 102 are connected via a wireless or wired network, and the terminal device 101 includes but is not limited to electronic devices such as desktop computers, mobile phones, mobile computers, tablet computers, media players, smart wearable devices, smart TVs, etc. The server 102 can be a single server, a server cluster consisting of several servers, or a cloud computing center. The server 102 can be an independent physical server, or a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0080] Each terminal device collects bill images and sends each bill image to the server for processing. The server applies the data set construction method in the embodiment of the present application to obtain the labeled element entity, which is called a data set. Of course, the data set construction process can also be completed by the terminal device.

[0081] Of course, the method provided in the embodiment of the present application is not limited to Figure 1 The application scenarios shown can also be used in other possible application scenarios, and the embodiments of the present application are not limited thereto. Figure 1 The functions that can be implemented by each device in the application scenario shown will be described in the subsequent method embodiments, and will not be described in detail here.

[0082] To further illustrate the technical solution provided by the embodiment of the present application, this is described in detail below in conjunction with the accompanying drawings and specific implementation methods. Although the embodiment of the present application provides the method operation steps shown in the following embodiments or drawings, more or fewer operation steps may be included in the method based on routine or no creative labor. In the steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided in the embodiment of the present application.

[0083] Combine the following Figure 1 The application scenario shown illustrates the technical solution provided by the embodiment of the present application. Figure 2 A schematic diagram of a bill style is shown. The bill includes multiple elements. When there are a large number of bills, it is necessary to automatically identify the style or type of the bill.

[0084] refer to Figure 3, the present application embodiment provides a method for constructing a data set, comprising the following steps:

[0085] S301, obtaining a bill image to be processed.

[0086] S302. For each bill image, perform OCR recognition on the bill image to determine the element information of each element entity in the bill image and the position information of each element entity; wherein the element information includes at least one of text information, table information and signature information.

[0087] S303: For each feature entity, determine the feature entity according to the feature information and location information.

[0088] S304: Based on each feature entity, label the feature entity using a label corresponding to the category of the feature entity.

[0089] S305: Determine a set of elements consisting of each annotated element entity in each bill image to be processed as a data set.

[0090] In the embodiment of the present application, after obtaining the bill image to be processed, OCR recognition is performed on each bill image to determine the element information of each element entity in the bill image and the position information of each element entity; in this way, the category of the corresponding element entity can be determined together with the element information and the position information, and then in the process of data annotation, the element entity can be annotated with the label corresponding to the category of the element entity, and the element set composed of the annotated element entities is a data set. The data set includes not only the element information, but also the position information, thus improving the accuracy of the bill recognition model obtained by using the data set as a sample for training, and improving the accuracy of bill recognition.

[0091] In S301, in order to obtain samples for model training, the bill images to be processed must first be obtained. In this process, a large number of bill images of different types can be obtained. If it is an electronic bill, the bill image can be directly captured; if it is a paper bill, the paper bill can be photographed to obtain the bill image.

[0092] In S302, for each bill image, OCR recognition technology is applied to identify the element information and position information of each element entity in each bill image. In the field of bills, element information includes at least one of text information, table information and signature information. Different bill types have different element entities. This is just an example and does not form a specific limitation. Specifically, the position information can be determined by the coordinates of the element information output by the OCR recognition technology.

[0093] Therefore, due to the different issuers of the bills, the bills often have very different formats, but the element entity information to be expressed is basically the same, and the locations where a large part of the element entities appear are also relatively fixed. In this way, the information of the text itself, the keywords around the text, and the location information of the text on the bill are combined. For such bills, most of the elements can be well covered, and the confidence of element extraction will also be increased during extraction.

[0094] Regarding S303, since a bill image includes multiple element entities, the category of each element entity can be determined based on the element information and position information.

[0095] Specifically, the category of the element entity can be determined by judging whether the element information includes preset keyword information. The preset keyword information represents the type of the element or data that can represent the element type. For example, the preset keyword information can be "date", "invoice issuer", etc., so the category of the element entity can be determined by the preset keyword.

[0096] In the first case, the element information includes preset keywords.

[0097] In this case, the preset keywords included in the feature information may be directly extracted to determine the category of the feature entity corresponding to the feature information.

[0098] In the second case, the element information does not include preset keywords.

[0099] In this case, if only the feature information is used, the category of the feature entity cannot be determined. Therefore, at this time, the adjacent feature entity of the feature entity can be determined based on the location information of the feature information, and the adjacent relationship can be determined by calculating the distance on the bill image. In a specific example, for example, if the current feature information does not include preset keyword information, the feature entity whose distance from the current feature entity on the bill image is less than the preset distance threshold is determined as the adjacent feature entity. In this way, the category of the current feature entity can be determined based on the preset keywords included in the adjacent feature entity.

[0100] Therefore, for the task of labeling bill information, OCR text recognition technology can well replace the step of manually inputting image text and extract all the location information and content information of the image text. The labeling data that retains the location information or adjacent element information will also be learned as a weight during training. In bill images with more complex information styles, especially when there is no keyword information around, the weight of the location information or adjacent element information will be very important. In the same type of bills, the specific location information of an element and the surrounding element layout information have a high similarity. If this information is used well, the corresponding element information can be better extracted in this type of image bill.

[0101] In S304, after the categories of each feature entity are determined, a label corresponding to the category of each feature entity is applied to mark the feature entity.

[0102] Specifically, for a bill image, the annotation process varies depending on the information that can be extracted. The annotation process includes the following situations:

[0103] The first case is when the category of the feature entity is a simple category.

[0104] The number of feature entities in the region where the feature entity of the simple category is located is one. Therefore, when the category of the feature entity is a simple category, the feature entity is directly marked as the label of the feature entity, for example, the label of the feature entity is "invoice" and the like.

[0105] For simple category elements, such as "keyword: element value" or "element value", they are directly marked as the label of the element, and an attribute entity is added to save the actual element value of the label.

[0106] In a specific example, Figure 4 In , an element text box (also called a span box) is a simple type of element, such as Figure 4 In the "Invoice" box, this span box only has "Invoice", and the element label is the bill title.

[0107] The second case is when the category of the feature entity is a composite category.

[0108] The number of element entities in the region where the element entity of the composite category is located is at least two. Therefore, when the category of the element entity is a composite category, a composite label is applied to label the current element entity. The composite label includes the labels of each element entity in the composite element and the corresponding element value.

[0109] For compound tags, such as "keyword 1: element value 1{keyword 2: element value 2...}" or "element value 1{element value 2...}", they need to be marked as compound tags, and the entity value stores the labels and element values ​​of each element, such as entity: "{label 1: element value 1}{label 2: element value 2}..."

[0110] In a specific example, still referring to Figure 4 "Bill issuer: XXX Billing date: 2021-06-01", there are two pieces of information in this span box, namely the bill issuer and the billing date. At this time, a composite tag COMPLEX needs to be defined to meet this situation.

[0111] For example, Figure 5 A schematic diagram of a labeling process is shown, wherein area 501 is the labeling of feature entities of a simple category, and area 502 is the labeling of feature entities of a complex category.

[0112] In addition, you can also use Excel to select labels and fill in element values ​​for the text extracted by OCR text recognition, such as Figure 6 .

[0113] It can also be combined with OCR recognition technology to add annotations to the OCR recognition results. That is, after the bill image is imported into the system, OCR text recognition is performed, and then the recognition results are displayed on the page. For the text recognition box on the page, the element label can be selected and the specific element value can be marked, such as Figure 7 Specifically, simple category elements can be labeled with text boxes only, and composite category elements can be labeled with labels and element values.

[0114] In summary, analogous to part-of-speech tagging in sequence tagging, element labels are analogous to part-of-speech tags, and element entities are analogous to words in sentences. By classifying and labeling the information on a bill through limited element labels, we can obtain a feature set contained in a label. The larger the set, the richer the feature information contained, and the more accurate the model learning. The simple category elements and the compound category elements are distinguished, and the compound category elements are specially labeled to improve the accuracy of the constructed data set.

[0115] The above labels can be replaced by the part of speech of the element information included in the element entity. Specifically, the part of speech tagging of sequence tagging can be applied to determine the label of each element entity. By labeling the element entity value with the element label, the part of speech tagging method can be used to determine the label to construct the data set.

[0116] In S305, after labeling each element entity in each bill image, a set of elements consisting of each labeled element entity in each bill image to be processed is determined as a data set. The data set can be used as a training sample to obtain a bill recognition model, and the bill recognition model is used to recognize the bill to be recognized, with a relatively high accuracy rate.

[0117] The present invention constructs a data set for extracting bill image information by combining OCR text recognition with NLP sequence annotation. This method can make full use of OCR recognition technology to extract text information and table information from bill images, and then classify and annotate the element entities in the bill by annotating the OCR recognition results. In this way, the position information and entity information of the text in the bill image can be retained at the same time, and the information that can be learned is more comprehensive and complete.

[0118] In the actual application process, in order to improve the annotation effect, before annotating each element entity, the recognition result is displayed through the page according to the preset display form, that is, each element entity is displayed. In this way, the preset display form can be referred to for annotation.

[0119] In a specific example, the preset display forms are different and the display methods are different.

[0120] If the preset display format is HTML format, displaying each element entity according to the preset display format includes: determining an element text box, and displaying the element text box in HTML format; wherein the element text box includes attribute information of the element to be displayed.

[0121] Among them, a continuous independent line of text is a text box, which can also be called a span box. A text box includes the attribute information of the element to be displayed, such as the font size, position and color of a line of text. Therefore, when the preset display format is HTML, the element text box is determined and displayed in HTML format.

[0122] In addition, if the element to be displayed is an element in a table, the attribute label of the element text box can also be determined, and the attribute label is displayed as the auxiliary attribute information of the element text box. The attribute label can be a table label, indicating that a row of text is in a table.

[0123] If the preset display format is json format, then displaying each element entity according to the preset display format includes: displaying the element key-value pair according to the determined correspondence between the element key-value pairs; wherein the key in the key-value pair is the serial number of the element to be displayed, and the value in the key-value pair is the nested json string of the element to be displayed.

[0124] Among them, Json (JavaScript Object Notation, JS object notation) is a lightweight data exchange format. When the preset display format is json format, the element key-value pair is displayed according to the correspondence relationship of the determined element key-value pair. Exemplarily, the key in the key-value pair is the serial number of the element to be displayed, and the value in the key-value pair is the nested json string of the element to be displayed.

[0125] For a specific example, see Figure 4 , showing a display effect diagram after identifying element entities, wherein area 401 is the HTML display part and area 402 is the JSON display part.

[0126] After applying multiple bill images to be processed to obtain a labeled data set, the data set is used as a training data set to train a bill recognition model. Figure 8 , showing a schematic diagram of a training method for a bill recognition model.

[0127] S801. Obtain a training data set, wherein the training data includes a data set obtained by applying any method of claims 1 to 6.

[0128] S802: Use the data set to train the pre-built neural network model until the neural network model converges to obtain a bill recognition model.

[0129] In the embodiment of the present application, the training data set is obtained by the data set construction method in the aforementioned embodiment. Therefore, the pre-constructed neural network model is trained using the data set until the neural network model converges to obtain a bill recognition model. The neural network model can be a convolutional neural network model, etc. Since the data set involved in the training is obtained by applying the data set construction method in the present application, the trained neural network model is relatively accurate, and when the neural network model is used to recognize bills, the recognition accuracy is high.

[0130] In addition, in the embodiment of the present application, the labeling process of the data set can also be performed in the following manner:

[0131] During the first annotation, pure manual annotation is required to provide a data set for initial model training; after the initial model is trained, the system automatically calls the bill extraction model for mechanical annotation. At the same time, in order to improve the accuracy of the annotation, the results of the mechanical annotation need to be manually corrected to correct the errors in the model annotation. Therefore, in addition to the first manual annotation, you can choose to merge the previous annotation data of the same type of bills with the same label, train them together, generate a new extraction model, and compare it with the previous extraction model, and select the extraction model with the best indicators to replace the bill extraction model. The bill extraction model is used for data annotation, which is different from the bill recognition model.

[0132] Therefore, by initially labeling a type of bill, a smaller set of labeled numbers can be obtained. Training on this set can obtain an initial model. This model can be used to mechanically label subsequent bills. Due to the small sample size, the model may be able to identify fewer element values, but it can greatly reduce the workload of manual labeling. In continuous iterative training, the effect of the model can be quickly improved, and in the process of manual correction, the accuracy of model recognition can be better determined, and the weak points of model training, that is, sparse elements, can be found in a targeted manner, and targeted labeling training can be performed. Moreover, since the labeling data consistency of similar bills with the same label is very high, automated training can be performed by simply replacing the data set; even similar bills with different labels can be combined for training by deleting labels or replacing label information.

[0133] This labeling process can well meet the needs of reducing the workload of data labeling. Through continuous iterative training and the use of mechanical labeling of intermediate models, the labeling work of simple labels can be reduced, and only complex labels and sparse elements can be labeled in a targeted manner, which is more targeted. At the same time, due to the consistency of each round of labeled data, in terms of training, only the data set needs to be replaced to train the model, which further simplifies the training cost of the model.

[0134] In summary, the embodiment of the present application extracts the text information and position information from the bill image through OCR recognition, that is, the content information of the text itself is read, and the position information of the text itself on the bill is retained. For element entities that have no obvious keywords and can only be judged by position or adjacent element information, the position information or adjacent element information is the key information. Usually, sequence annotations will not mark the position of the text on the image when annotating, so it will be difficult to extract such elements. After combining OCR recognition, this key position information and related information of adjacent elements can be retained.

[0135] like Fig. 9As shown, based on the same inventive concept as the above-mentioned method for constructing a data set, an embodiment of the present application also provides a data set construction device, which includes an image acquisition module 91, an image recognition module 92, a category determination module 93, a labeling module 94 and a data set determination module 95.

[0136] An image acquisition module 91 is used to acquire the image of the bill to be processed;

[0137] The image recognition module 92 is used to perform OCR recognition on each bill image to determine the element information of each element entity in the bill image and the position information of each element entity; wherein the element information includes at least one of text information, table information and signature information;

[0138] A category determination module 93, for determining the category of each element entity according to the element information and the location information;

[0139] A labeling module 94 is used to label each feature entity by applying a label corresponding to the category of the feature entity;

[0140] The data set determination module 95 is used to determine that the element set consisting of each annotated element entity in each bill image to be processed is a data set.

[0141] In some exemplary embodiments, the category determination module 93 is specifically configured to:

[0142] If the element information does not include preset keyword information, the adjacent element entity of the element entity is determined according to the location information of the element information, and the category of the element entity is determined according to the adjacent element entity; wherein the distance between the adjacent element entity and the element entity on the bill image is less than the preset distance threshold.

[0143] In some exemplary embodiments, the annotation module 94 is specifically used to:

[0144] If the category of the feature entity is a simple category, the feature entity is marked as the label of the feature entity; wherein the number of feature entities in the region where the feature entity of the simple category is located is one;

[0145] If the category of the feature entity is a composite category, the feature entity is marked as a composite label; wherein the composite label includes the labels of each feature entity in the composite feature and the corresponding feature value; the number of feature entities in the area where the composite category feature entity is located is at least two.

[0146] In some exemplary embodiments, a display module is further included, and the display module is used to, based on the category of each feature entity, apply a label feature entity corresponding to the category before labeling the feature entity:

[0147] Display each element entity according to the preset display form.

[0148] In some exemplary embodiments, if the preset display format is HTML format, the display module is specifically used to:

[0149] Determine an element text box, and display the element text box in HTML format; wherein the element text box includes attribute information of the element to be displayed;

[0150] If the default display format is json, the display module is specifically used for:

[0151] The feature key-value pairs are displayed according to the determined correspondence between the feature key-value pairs; wherein the key in the key-value pair is the serial number of the feature to be displayed, and the value in the key-value pair is the nested JSON string of the feature to be displayed.

[0152] In some exemplary embodiments, the display module further includes, before determining the element text box and displaying the element text box in HTML format:

[0153] If the element to be displayed is an element in a table, determine the attribute label of the element text box;

[0154] The attribute label is displayed as auxiliary display information of the feature text box.

[0155] The apparatus for constructing a data set in the embodiment of the present application and the method for constructing a data set described above adopt the same inventive concept and can achieve the same beneficial effects, which will not be described in detail herein.

[0156] like Fig.10 As shown, based on the same inventive concept as the above-mentioned bill recognition model training method, the embodiment of the present application also provides a bill recognition model training construction device, which includes a data set acquisition module 1001 and a training module 1002.

[0157] A data set acquisition module 1001 is used to acquire a training data set, wherein the training data includes a data set obtained by the method of the second aspect above;

[0158] The training module 1002 is used to train the pre-built neural network model using the data set until the neural network model converges to obtain a bill recognition model.

[0159] The training device for the bill recognition model proposed in the embodiment of the present application and the training method for the bill recognition model mentioned above adopt the same inventive concept and can achieve the same beneficial effects, which will not be repeated here.

[0160] Based on the same inventive concept as the above-mentioned method for constructing a data set, the present application embodiment also provides an electronic device, which can be a desktop computer, a portable computer, a smart phone, a tablet computer, a personal digital assistant (PDA), a server, etc. Fig.11 As shown, the electronic device may include a processor 111 and a memory 112 .

[0161] The processor 111 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, and may implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.

[0162] The memory 112 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 112 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.

[0163] Based on the same inventive concept as the above-mentioned method for constructing a data set, the present application embodiment also provides an electronic device, which can be a desktop computer, a portable computer, a smart phone, a tablet computer, a personal digital assistant (PDA), a server, etc. Fig.12 As shown, the electronic device may include a processor 121 and a memory 122 .

[0164] The processor 121 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, and may implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.

[0165] The memory 122 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 122 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.

[0166] A person of ordinary skill in the art can understand that: all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiments; the above computer storage medium can be any available medium or data storage device that can be accessed by a computer, including but not limited to: mobile storage devices, random access memory (RAM, Random Access Memory), magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO)), optical storage (such as CD, DVD, BD, HVD, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)) and other media that can store program codes.

[0167] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of each embodiment of the present application. The aforementioned storage medium includes: a mobile storage device, a random access memory (RAM, Random Access Memory), a magnetic storage device (such as a floppy disk, a hard disk, a magnetic tape, a magneto-optical disk (MO), etc.), an optical storage device (such as CD, DVD, BD, HVD, etc.), and a semiconductor memory (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), a solid-state drive (SSD)) and other various media that can store program codes.

[0168] In some possible implementations, various aspects of the method provided by the present disclosure may also be implemented in the form of a program product, which includes a program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of the method according to various exemplary embodiments of the present disclosure described above in this specification. For example, the computer device may execute the transaction information processing method recorded in the embodiment of the present disclosure. The program product may adopt any combination of one or more readable media.

[0169] In the technical solution of this application, the collection, dissemination, and use of data are in compliance with the requirements of relevant national laws and regulations.

[0170] The above embodiments are only used to introduce the technical solutions of the present application in detail, but the description of the above embodiments is only used to help understand the methods of the embodiments of the present application and should not be understood as limiting the embodiments of the present application. Any changes or substitutions that can be easily thought of by technicians in this technical field should be included in the protection scope of the embodiments of the present application.

Claims

1. A method for constructing a data set, characterized in that: include: Get the bill image to be processed; For each of the bill images, perform OCR recognition on the bill image to determine the element information of each element entity in the bill image and the position information of each element entity; wherein the element information includes at least one of text information, table information and signature information; For each element entity, determining a category of the element entity according to the element information and the location information; If the element to be displayed is an element in a table, an attribute label of the element text box is determined, and the attribute label is displayed as auxiliary display information of the element text box; If the preset display format is HTML, determine the element text box and display the element text box in HTML format; if the preset display format is JSON, display the element key-value pair according to the determined correspondence between the element key-value pairs; wherein the element text box includes the attribute information of the element to be displayed; the key in the key-value pair is the serial number of the element to be displayed, and the value in the key-value pair is the nested JSON string of the element to be displayed; Based on each feature entity, label the feature entity using a label corresponding to the category of the feature entity; Determine a set of elements formed by each annotated element entity in each of the bill images to be processed as a data set; The applying a label corresponding to the category of the feature entity to mark the feature entity includes: If the labeling operation for the element entity is not the first labeling in the entire labeling process, the element entity is labeled using the label corresponding to the category of the element entity through the bill extraction model; wherein the bill extraction model is updated using the last labeling result.

2. The method according to claim 1, characterized in that The determining the category of the feature entity according to the feature information and the location information includes: If the element information does not include preset keyword information, the adjacent element entity of the element entity is determined according to the location information of the element information, and the category of the element entity is determined according to the adjacent element entity; wherein the distance between the adjacent element entity and the element entity on the bill image is less than a preset distance threshold.

3. The method according to claim 1, characterized in that Based on the category of each feature entity, labeling the feature entity with a label corresponding to the category includes: If the category of the feature entity is a simple category, the feature entity is marked as a label of the feature entity; wherein the number of feature entities in the area where the feature entity of the simple category is located is one; If the category of the feature entity is a composite category, the feature entity is marked as a composite label; wherein the composite label includes labels of each feature entity in the composite feature and corresponding feature values; and the number of feature entities in the area where the feature entity of the composite category is located is at least two.

4. A training method for a bill recognition model, characterized in that: include: Obtaining a training data set, wherein the training data includes a data set obtained by applying any of the methods described in claims 1 to 3; The pre-built neural network model is trained using the data set until the neural network model converges to obtain a bill recognition model.

5. A device for constructing a data set, characterized in that: include: An image acquisition module, used to acquire the image of the bill to be processed; An image recognition module, for performing OCR recognition on each of the bill images, and determining element information of each element entity in the bill image and position information of each element entity; wherein the element information includes at least one of text information, table information and signature information; A category determination module, used for determining the category of each element entity according to the element information and the location information; A display module, used for: if the element to be displayed is an element in a table, determining an attribute label of an element text box, and displaying the attribute label as attached display information of the element text box; The display module is further used to: if the preset display format is HTML, determine the element text box and display the element text box in HTML format; if the preset display format is JSON, display the element key-value pair according to the determined correspondence relationship of the element key-value pair; wherein the element text box includes the attribute information of the element to be displayed; the key in the key-value pair is the serial number of the element to be displayed, and the value in the key-value pair is the nested JSON string of the element to be displayed; A labeling module, used for labeling each feature entity by applying a label corresponding to the category of the feature entity; A data set determination module, used to determine a set of elements consisting of each annotated element entity in each of the bill images to be processed as a data set; The labeling module is specifically used for: if the labeling operation on the element entity is not the first labeling in the entire labeling process, the element entity is labeled by using the label corresponding to the category of the element entity through the bill extraction model; wherein, the bill extraction model is updated using the last labeling result.

6. The device according to claim 5, characterized in that The category determination module is specifically used for: If the element information does not include preset keyword information, the adjacent element entity of the element entity is determined according to the location information of the element information, and the category of the element entity is determined according to the adjacent element entity; wherein the distance between the adjacent element entity and the element entity on the bill image is less than a preset distance threshold.

7. The device according to claim 5, characterized in that The annotation module is specifically used for: If the category of the feature entity is a simple category, the feature entity is marked as a label of the feature entity; wherein the number of feature entities in the area where the feature entity of the simple category is located is one; If the category of the feature entity is a composite category, the feature entity is marked as a composite label; wherein the composite label includes labels of each feature entity in the composite feature and corresponding feature values; and the number of feature entities in the area where the feature entity of the composite category is located is at least two.

8. A training device for a bill recognition model, characterized in that: include: A data set acquisition module, used to acquire a training data set, wherein the training data includes a data set obtained by applying any method described in claims 1 to 3; The training module is used to train the pre-built neural network model using the data set until the neural network model converges to obtain a bill recognition model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Method and device for establishing bill type identification model and method and device for identifying bill type

    CN113033534A