Form processing method, apparatus, device, and readable storage medium
By preprocessing business documents and parsing them using a deep instance segmentation model, combined with a form language processing model, the problem of low processing efficiency for different types of documents is solved, achieving fast and accurate information extraction and format unification.
Patent Information
- Application Number
- CN202211329329.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-27
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2042-10-27
AI Technical Summary
Existing technologies cannot efficiently process different types of business documents, and are prone to data confusion, errors, or omissions.
By preprocessing the business documents to be processed, including data cleaning and format conversion, and using deep instance segmentation model and form language processing model for parsing and information extraction, the target information can be obtained.
It enables fast and accurate processing of different types of business documents, ensuring information integrity and format consistency, and improving processing efficiency and accuracy.
Smart Images

Figure CN115688716B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of information processing, and in particular to a form processing method and device, equipment and a readable storage medium. BACKGROUND
[0002] With the development of economy, more and more enterprises use electronic forms to manage their business data, so in order to meet business needs, many electronic invoices will be generated accordingly when business transactions are carried out. These electronic invoices of different types of business reflect the business information of the enterprise.
[0003] In order to better manage a variety of electronic invoices, it is necessary to analyze different types of invoices to obtain relevant information of the electronic invoices and enter the relevant information of the electronic invoices into an information business management system. However, the existing electronic invoice processing method can only analyze and process a certain type of business electronic invoice, cannot efficiently process different types of business invoices, and when extracting relevant information of electronic invoices from a variety of different formats of electronic invoices, it is very easy to cause confusion, resulting in extraction data errors or omissions. SUMMARY
[0004] The present application aims to at least solve one of the above technical defects, and in view of this, the present application provides a form processing method, device, equipment and readable storage medium, which can efficiently process different types of business invoices.
[0005] A form processing method, comprising:
[0006] Obtaining a business invoice to be processed;
[0007] Preprocessing the business invoice to be processed to obtain a target business invoice;
[0008] Analyzing the target business invoice to obtain a first analysis result of the target business invoice;
[0009] Inputting the first analysis result of the target business invoice into a preset deep instance segmentation model for processing to obtain a second analysis result of the target business invoice;
[0010] Analyzing the second analysis result of the target business invoice to obtain first target information of the target business invoice;
[0011] Inputting the first target information of the target business invoice into a preset form language processing model for processing to obtain second target information of the target business invoice.
[0012] Preferably, the preprocessing of the business invoice to be processed to obtain a target business invoice comprises:
[0013] performing data cleaning on the to-be-processed business document, and removing a business document with incorrect or incomplete content from the to-be-processed business document to obtain a second to-be-processed business document;
[0014] identifying a format of the second to-be-processed business document;
[0015] translating the format of the second to-be-processed business document into a format meeting preset format requirements according to the preset format requirements to obtain the target business document.
[0016] Preferably, the parsing of the target business document to obtain a first analysis result of the target business document comprises:
[0017] judging whether the target business document contains a document or picture that can be parsed or a document or picture that cannot be parsed;
[0018] if the target business document contains a document or picture that can be parsed, parsing the document or picture that can be parsed in the target business document, and converting the parsed document into a picture to obtain a first analysis result of the target business document;
[0019] if the target business document contains a document or picture that cannot be parsed, performing content recognition on the document or picture that cannot be parsed in the target business document, and converting the parsed document into a picture to obtain a second analysis result of the target business document;
[0020] taking the first analysis result of the target business document and the second analysis result of the target business document as the first analysis result of the target business document.
[0021] Preferably, the parsing of the target business document to obtain a first analysis result of the target business document comprises:
[0022] parsing the document or picture that can be parsed in the target business document to obtain a parsing result of each character of the target business document, coordinate information of each character of the target business document, and a text display attribute of each character of the target business document;
[0023] converting the parsed document into a picture to obtain a picture corresponding to the target business document;
[0024] The image features of the picture corresponding to the target business document, the recognition result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document are taken as the first analysis result of the target business document.
[0025] Preferably, the content recognition is performed on the document or picture that cannot be parsed in the target business document, and the document that cannot be parsed is converted into a picture to obtain a second analysis result of the target business document, including:
[0026] The content recognition is performed on the document or picture that cannot be parsed in the target business document to obtain the recognition result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document.
[0027] The document that cannot be parsed is converted into a picture to obtain a picture corresponding to the target business document.
[0028] The image features of the picture corresponding to the target business document, the recognition result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document are taken as the second analysis result of the target business document.
[0029] Preferably, the creating process of the preset deep instance segmentation model includes:
[0030] The analysis result corresponding to a training business document is taken as a training sample, target information contained in the analysis result corresponding to the training business document is taken as a sample label, and a deep instance segmentation model is trained and obtained.
[0031] The analysis result corresponding to the training business document includes image features of a picture corresponding to the training business document, recognition results and identification results of each character of the training business document, coordinate information of each character of the training business document in the training business document, and text display attributes of each character of the training business document.
[0032] The target information contained in the analysis result corresponding to the training business document includes character classification of each character of the training business document, each business field name corresponding to the training business document, values corresponding to each business field of the training business document, content of descriptive words corresponding to the training business document, and a border of each business field corresponding to the training business document.
[0033] The target information contained in the analysis result corresponding to the training business document includes character classification of each character of the training business document, each business field name corresponding to the training business document, values corresponding to each business field of the training business document, content of descriptive words corresponding to the training business document, and a border of each business field corresponding to the training business document.
[0034] Preferably, the first target information of the target business document includes the name of each business field corresponding to the target business document, and the first target information of the target business document is input into a preset form language processing model for processing to obtain second target information of the target business document, which includes:
[0035] The name of each business field corresponding to the target business document in the first target information of the target business document is extracted to obtain the name of each business field corresponding to the target business document.
[0036] The name of each business field corresponding to the target business document is input into a preset form language processing model for business standard field alignment processing to obtain the value corresponding to each business field in the target business document.
[0037] Preferably, the creation process of the preset form language processing model includes:
[0038] The first target information corresponding to the training business document is taken as a training sample, and the second target information contained in the first target information corresponding to the training business document is taken as a sample label, and a form language processing model is trained and obtained.
[0039] The first target information corresponding to the training business document includes the analysis result and the recognition result of each character of the training business document corresponding to the training business document.
[0040] The second target information contained in the first target information corresponding to the training business document includes the value corresponding to each business field corresponding to the analysis result and the recognition result of each character of the training business document corresponding to the training business document, and the content of the descriptive text corresponding to the training business document.
[0041] The second target information contained in the first target information corresponding to the training business document includes the value corresponding to each business field corresponding to the analysis result and the recognition result of each character of the training business document corresponding to the training business document, and the content of the descriptive text corresponding to the training business document.
[0042] A form processing device includes:
[0043] A data acquisition unit is configured to acquire a business document to be processed.
[0044] A data preprocessing unit is configured to preprocess the business document to be processed to obtain a target business document.
[0045] A data analysis unit is configured to analyze the target business document to obtain a first analysis result of the target business document.
[0046] A first data processing unit is configured to input the first analysis result of the target business document into a preset deep instance segmentation model for processing to obtain a second analysis result of the target business document.
[0047] a second data processing unit configured to analyze a second parsing result of the target business document to obtain first target information of the target business document;
[0048] a third data processing unit configured to input the first target information of the target business document into a preset form language processing model to obtain second target information of the target business document.
[0049] A form processing device, comprising: one or more processors, and a memory;
[0050] The memory stores computer readable instructions, which, when executed by the one or more processors, implement the steps of the form processing method according to any one of the preceding embodiments.
[0051] A readable storage medium, the readable storage medium stores computer readable instructions, which, when executed by one or more processors, cause the one or more processors to implement the steps of the form processing method according to any one of the preceding embodiments.
[0052] As can be seen from the above technical solutions, when a large number of different types of business documents need to be processed quickly, the method provided by the embodiments of the present application can obtain a to-be-processed business document; the to-be-processed business document is preprocessed so that the format of the to-be-processed business document can be converted into a preset format, thereby obtaining a target business document; after obtaining the target business document, the target business document is parsed to obtain a first parsing result of the target business document, which includes all information of the target business document; by analyzing the first parsing result of the target business document, all information of the target business document can be understood; therefore, after obtaining the first parsing result of the target business document, the first parsing result of the target business document is input into a preset deep instance segmentation model to obtain a second parsing result of the target business document, wherein the first parsing result of the target business document is analyzed by using the preset deep instance segmentation model to understand all information of the target business document; after obtaining the second parsing result of the target business document, the second parsing result of the target business document is further analyzed to obtain first target information of the target business document; the first target information of the target business document is input into a preset form language processing model to obtain second target information of the target business document.
[0053] The method provided by the embodiment of the application can preprocess business bills of different formats, and analyze some important note information and unexpected situation description information in the business bills, thereby analyzing all related information of the business bills, helping to extract target information from all information of the business bills, and further processing the target information by using a language processing model, so that the form information can be quickly obtained, and the obtained target information meets preset format requirements. The method provided by the embodiment of the application can be applied to form processing in different scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0055] Figure 1 A flowchart of a form processing method provided by the embodiment of the present application;
[0056] Figure 2 A structure schematic diagram of a form processing device provided by the embodiment of the present application;
[0057] Figure 3 A hardware structure block diagram of a form processing device disclosed by the embodiment of the present application. DETAILED DESCRIPTION
[0058] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0059] When it is necessary to quickly process a large number of business bills of different types, since most of the current form processing schemes are difficult to adapt to complex and changeable business requirements, the present applicant has researched a form processing scheme. The form processing method preprocesses business bills of different formats, and analyzes some important note information and unexpected situation description information in the business bills, thereby analyzing all information of the business bills, helping to extract target information from all information of the business bills, and further processing the target information by using a language processing model, so that the form information can be quickly obtained, and the obtained target information meets preset format requirements, and the method can be applied to form processing in different scenarios.
[0060] The method provided by the embodiments of the present application can be used in many general or specific computing device environments or configurations. For example, personal computers, server computers, handheld or laptop devices, tablet devices, multiprocessor systems, distributed computing environments that include any of the above systems or devices, and the like.
[0061] The form processing method provided by the embodiments of the present application can be applied to various business management systems, and can also be applied to various computer terminals or intelligent terminals. The execution subject of the form processing method can be a processor of a computer terminal or an intelligent terminal or a server.
[0062] The form processing method provided by the embodiments of the present application will be described below in combination with Figure 1 The flow of the form processing method provided by the embodiments of the present application is shown in FIG. 1, which can include the following steps: Figure 1
[0063] In step S101, a to-be-processed business document is obtained.
[0064] Specifically, in actual application, with the development of economy, more and more enterprises or units use electronic forms to manage their business data. Therefore, in order to meet the business needs, many different types of electronic documents are generated when the business transactions are performed. The electronic documents of different types of businesses reflect the business information of the enterprises.
[0065] In order to better manage the various types of electronic documents, the information of the electronic documents is obtained by analyzing different single documents, and the related information of the electronic documents is input into the information business management system.
[0066] Therefore, when the business data of an enterprise or a unit needs to be managed, the to-be-processed business document of the unit or the enterprise can be obtained, so that the business situation of the unit or the enterprise can be understood by analyzing the to-be-processed business document of the unit or the enterprise.
[0067] In step S102, the to-be-processed business document is preprocessed to obtain a target business document.
[0068] Specifically, in actual application, the business types of different enterprises or units are different, and even the same enterprise or the same unit can have various types of businesses. Different businesses can include a variety of business documents. In order to more quickly process different types of business documents, the to-be-processed business document can be preprocessed to obtain a target business document.
[0069] Different types of business documents can have many different formats of business documents. The to-be-processed business document is preprocessed so that the format of the to-be-processed business document can be uniformly converted to obtain a target business document with complete data and clear content and uniform format.
[0070] The preprocessing of the to-be-processed business document can improve the parsing speed of the to-be-processed business document and reduce the omission or error of the data of the to-be-processed business document.
[0071] In step S103, the target business document is parsed to obtain a first parsing result of the target business document.
[0072] Specifically, as described above, the method provided in the embodiments of the present application can preprocess the to-be-processed business document, so that the target business document can be obtained.
[0073] The target business document is a business document to be parsed. The relevant information of the target business document can be obtained by parsing the target business document.
[0074] Therefore, after the target business document is obtained, the target business document can be parsed to obtain a first parsing result of the target business document.
[0075] The first parsing result of the target business document includes all information of the target business document.
[0076] For example,
[0077] Currently, the information of most types of business documents is embodied in a format similar to key-value-content, which represents the relevant information of the business document.
[0078] Among them,
[0079] The key can represent the name of the business field of the business document;
[0080] The value can represent the value corresponding to the business field in the business document;
[0081] The content can represent the descriptive text in the business document, such as remarks and postscripts.
[0082] For example, the key-value of a certain business document is "name-Zhang San", which means that the name of the business field of the business document is "name" and the value corresponding to the business field is "Zhang San".
[0083] If the content of a certain business document is "remark - Zhang San is a student", the descriptive text of the business document can be "Zhang San is a student".
[0084] In addition, since the format of the business document is a table, in general, the related information of the business document can also include the Bounding Box corresponding to each business field of the business document. Bounding Box is a professional term in computer vision, which refers to a rectangular frame that can frame a certain image element in an image. This rectangular frame is generally obtained by fitting coordinates using deep learning training. The purpose of setting Bounding Box in the business document is to extract the frame corresponding to each business field in the business document.
[0085] Therefore, by analyzing the target business document, all information of the business document can be understood.
[0086] In step S104, the first analysis result of the target business document is input into a preset deep instance segmentation model for processing to obtain a second analysis result of the target business document.
[0087] Specifically, as described above, the method provided by the embodiments of the present application can analyze the target business document to obtain the first analysis result of the target business document, and the first analysis result includes all information of the target business document.
[0088] If the first analysis result of the target business document is directly analyzed to obtain the target information, it may be relatively time-consuming and the processing efficiency is not high.
[0089] Therefore, in order to further obtain the target information from all information of the target business document, after obtaining the first analysis result of the target business document, the first analysis result of the target business document can be input into a preset deep instance segmentation model for processing, so that the second analysis result of the target business document can be obtained.
[0090] Compared with the first analysis result of the target business document, the second analysis result of the target business document includes less information of the target business document than the first analysis result.
[0091] Therefore, the processing speed of analyzing the second analysis result to obtain the target information is higher than that of analyzing the first analysis result to obtain the target information.
[0092] Among them,
[0093] The creation process of the preset deep instance segmentation model can include the following:
[0094] The target information contained in the analysis result corresponding to the training business document is taken as a sample label, and training is performed to obtain;
[0095] The target information contained in the analysis result corresponding to the training business document can include:
[0096] The target information contained in the analysis result corresponding to the training business document can include:
[0097] The image features of the picture corresponding to the training business document;
[0098] The analysis result and the recognition result of each character of the training business document;
[0099] The coordinate information of each character of the training business document in the training business document;
[0100] The text display attribute of each character of the training business document;
[0101] The target information contained in the analysis result corresponding to the training business document can include:
[0102] The character classification of each character of the training business document and each business field name corresponding to the training business document;
[0103] The value corresponding to each business field corresponding to the training business document;
[0104] The content of the descriptive text corresponding to the training business document;
[0105] The border of each business field corresponding to the training business document.
[0106] Step S105, analyzing the second analysis result of the target business document to obtain the first target information of the target business document.
[0107] Specifically, as known from the above description, the method provided in the embodiments of the present application can input the first analysis result of the target business document into a preset deep instance segmentation model for processing to obtain the second analysis result of the target business document.
[0108] The second analysis result includes the related information of the target business document, and the related information of the target business document included in the second analysis structure includes the target information that is intended to be obtained.
[0109] The second analysis result is an analysis result obtained through secondary analysis.
[0110] The processing speed of analyzing the second analysis result is relatively efficient compared with the processing speed of analyzing the first analysis result.
[0111] Therefore, after obtaining the second analysis result, the second analysis result of the target business document can be further analyzed to obtain the first target information of the target business document.
[0112] As introduced above, the preset deep instance segmentation model can extract the business field name of the target business document, the value corresponding to each business field, and the edge frame of the region block corresponding to the descriptive text in the target business document.
[0113] In addition, the region coordinates where the edge frame of the region block corresponding to each business field text in the target business document is located, i.e., the edge frame (Boundingbox) corresponding to each business field, can also be obtained.
[0114] Through the edge frame corresponding to each business field of the target business document, the accompanying structural information of the target business document can be understood, which refers to the relative positional relationship between the edge frames of the region blocks corresponding to the texts obtained by the positions of the edge frames of the region blocks corresponding to all texts in the whole document.
[0115] For example,
[0116] The edge frame x of the region block corresponding to the text of a certain business field and the edge frame y of the region block corresponding to the text of a certain business field are in a left-right adjacent relationship, while the edge frame z of the region block corresponding to the text of a certain business field and the edge frame x of the region block corresponding to the text of a certain business field are in an up-down adjacent relationship.
[0117] Therefore, the first target information can include the business field name of the target business document, the value corresponding to each business field, and the edge frame of the region block corresponding to the descriptive text in the target business document.
[0118] In step S106, the first target information of the target business document is input into a preset form language processing model for processing to obtain the second target information of the target business document.
[0119] Specifically, as introduced above, the method provided by the embodiments of the present application can analyze the second analysis result of the target business document to obtain the first target information of the target business document.
[0120] The first target information can include the business field name of the target business document, the value corresponding to each business field, and the edge frame of the region block corresponding to the descriptive text in the target business document.
[0121] In actual application, the first target information of the target business document can be more and disordered. In order to obtain the target information meeting the preset format requirement more accurately, after obtaining the first target information of the target business document, the first target information of the target business document can be further input into a preset form language processing model for processing to obtain second target information of the target business document.
[0122] wherein,
[0123] The creation process of the preset form language processing model can include the following:
[0124] The first target information corresponding to the training business document can be taken as a training sample, and the second target information contained in the first target information corresponding to the training business document can be taken as a sample label, and the form language processing model can be trained and obtained.
[0125] wherein,
[0126] The first target information corresponding to the training business document can include the analysis result and the recognition result of each character of the training business document corresponding to the training business document.
[0127] The second target information contained in the first target information corresponding to the training business document can include the value corresponding to each business field corresponding to the analysis result and the recognition result of each character of the training business document corresponding to the training business document, and the content of the descriptive text corresponding to the training business document.
[0128] For example,
[0129] After obtaining the first target information, the method provided in the embodiment of the application can input the first target information of the target business document into a preset form language processing model for processing to obtain second target information of the target business document. The second target information can include a target business field name in the target business document meeting a preset format requirement and a value corresponding thereto.
[0130] As can be seen from the above technical solution, when a large number of different types of business documents need to be processed quickly, the method provided in the embodiment of the application can preprocess business documents of different formats, and analyze some important note information and unexpected situation description information in the business documents. Thus, all information of the business document can be analyzed, which is helpful to extract target information from all information of the business document, and the language processing model can be used to further process the target information. The method not only can quickly process forms to obtain form information, but also can make the obtained target information meet the preset format requirement. The method provided in the embodiment of the application can be applied to form processing in different scenarios.
[0131] From the above introduction, the method provided in the embodiments of the present application can preprocess the to-be-processed business document to obtain a target business document. Next, the process is introduced. The process can include the following steps.
[0132] In step S201, data cleaning is performed on the to-be-processed business document, and a business document with content errors or incomplete content is removed to obtain a second to-be-processed business document.
[0133] Specifically, as introduced above, the method provided in the embodiments of the present application can obtain a to-be-processed business document.
[0134] In actual application, because the business types of enterprises or units are various, and the generated business documents are also various, some business documents with filling errors or incomplete content can be generated.
[0135] For example, in the to-be-processed business document, there can be some business documents with format errors or data errors or missing some necessary information.
[0136] After obtaining the to-be-processed business document, if the to-be-processed business document is directly parsed, the processing speed can be relatively low in efficiency.
[0137] Therefore, after obtaining the to-be-processed business data, the to-be-processed business document can be further subjected to data cleaning, and a business document with content errors or incomplete content is removed to obtain a second to-be-processed business document.
[0138] The second to-be-processed business document can be data left after removing the to-be-processed business document with errors or incomplete content.
[0139] In step S202, the format of the second to-be-processed business document is identified.
[0140] Specifically, as introduced above, the method provided in the embodiments of the present application can perform data cleaning on the to-be-processed business document, and remove a business document with content errors or incomplete content to obtain a second to-be-processed business document.
[0141] The second to-be-processed business document is a business document that needs to be further subjected to parsing processing to obtain the target information.
[0142] Therefore, by parsing the second to-be-processed business document, the target information can be obtained.
[0143] In order to improve the speed of resolving the second to-be-processed business document, the format of the second to-be-processed business document can be identified first, so that unified format processing can be performed according to different format requirements of the business document.
[0144] In step S203, the format of the second to-be-processed business document is converted into a format meeting the preset format requirement according to the preset format requirement, so as to obtain the target business document.
[0145] Specifically, as described above, the method provided in the embodiments of the present application can identify the format of the second to-be-processed business document, so that the format of the second to-be-processed business document can be obtained.
[0146] Since the processing mode and speed of different formats are different, in order to improve the processing speed of the second to-be-processed business document, the format of the second to-be-processed business document can be converted into a format meeting the preset format requirement according to the preset format requirement, so that the target business document can be obtained.
[0147] The format of the target business document is a unified format, and when the target business document needs to be resolved, the processing speed of the target business document can be effectively improved.
[0148] As can be seen from the above technical solutions, the method provided in the embodiments of the present application can preprocess the to-be-processed business document, delete erroneous data or data with incomplete content in the to-be-processed business document, and perform unified format processing on the business data left after the erroneous data or data with incomplete content are deleted, so that the target business document can be obtained. The format of the target business document is a unified format, and when the target business document needs to be resolved, the processing speed of the target business document can be effectively improved.
[0149] As described above, the method provided in the embodiments of the present application can resolve the target business document to obtain a first resolution result of the target business document. Next, the process will be introduced, which can include the following steps:
[0150] In step S301, it is judged whether the target business document has a document or picture that can be resolved or a document or picture that cannot be resolved.
[0151] Specifically, as described above, after the to-be-processed business document is preprocessed, the method provided in the embodiments of the present application can obtain the target business data with a unified format.
[0152] Since the to-be-processed business document has been formatted, the format of the target business data is a unified format.
[0153] The target business document can include a document and can also include a picture.
[0154] The corresponding analysis manner is different for different documents and pictures.
[0155] Therefore, after obtaining the target business document, it can be judged whether the target business document has a document or a picture that can be analyzed or a document or a picture that cannot be analyzed.
[0156] If the target business document has a document or a picture that can be analyzed, step S302 can be performed to analyze the document or the picture that can be analyzed in the target business document.
[0157] If the target business document has a document or a picture that cannot be analyzed, step S303 can be performed to analyze the document or the picture that cannot be analyzed in the target business document.
[0158] Step S302 analyzes the document or the picture that can be analyzed in the target business document, and converts the document that can be analyzed into a picture to obtain a first analysis result of the target business document.
[0159] Specifically, as known from the above description, the method provided in the embodiments of the present application can determine whether the target business document has a document or a picture that can be analyzed or a document or a picture that cannot be analyzed. When the target business document has a document or a picture that can be analyzed, it indicates that the target business document has a document or a picture that can be analyzed, which can be obtained by analyzing the target information included in the document or the picture that can be analyzed in the target business document.
[0160] Therefore, after determining that the target business document has a document or a picture that can be analyzed, the document or the picture that can be analyzed in the target business document can be analyzed, and the document that can be analyzed is converted into a picture to obtain a first analysis result of the target business document.
[0161] The document that can be analyzed is converted into a picture, which is helpful for batch processing by using the preset deep instance segmentation model to obtain the target information included in the document or the picture that can be analyzed in the target business document.
[0162] Step S303 analyzes the content of the document or the picture that cannot be analyzed in the target business document, and converts the document that cannot be analyzed into a picture to obtain a second analysis result of the target business document.
[0163] Specifically, as can be known from the above description, the method provided in the embodiments of the present application can determine whether the target business document contains a document or picture that can be parsed or a document or picture that cannot be parsed. When the target business document contains a document or picture that cannot be parsed, it indicates that the document or picture that can be parsed in the target business document cannot be processed by parsing, and the document or picture that cannot be parsed in the target business document needs to be processed by recognition to obtain the target information included in the document or picture that cannot be parsed in the target business document.
[0164] Therefore, after determining that the target business document contains a document or picture that cannot be parsed, the document or picture that cannot be parsed in the target business document can be processed by recognition, and the document that cannot be parsed is converted into a picture to obtain a second analysis result of the target business document.
[0165] The conversion of the document that cannot be parsed into a picture helps to use the preset deep instance segmentation model to perform batch processing to obtain the target information included in the document or picture that cannot be parsed in the target business document.
[0166] For example, after determining that the target business document contains a document or picture that cannot be parsed, the document or picture that cannot be parsed in the target business document can be processed by OCR recognition, so that the recognition result of each character in the target business document, the position of each character in the target business document, and the font size of each character can be obtained.
[0167] The recognition result of each character in the target business document, the position of each character in the target business document, and the font size of each character can be obtained.
[0168] OCR is the abbreviation of Optical Character Recognition, and OCR (Optical Character Recognition) refers to the process in which an electronic device (such as a scanner or a digital camera) checks characters printed on paper, determines their shapes by detecting light and dark patterns, and then translates the shapes into computer text by using character recognition methods. In short, the OCR recognition technology is to use OCR recognition software to extract and recognize the characters in the picture by using the OCR software, and convert them into searchable data.
[0169] The OCR recognition technology can be applied to the following application scenarios:
[0170] (1) Certificate OCR recognition;
[0171] Certificate OCR recognition technology was initially based on PC, and has started to develop towards mobile terminals in recent years. At present, mature ones include identity card recognition, driving license recognition, driving license recognition, passport recognition, etc.
[0172] (2) Bank card OCR recognition;
[0173] Bank card OCR recognition is mainly used for mobile payment card binding, which is a very technical sub-OCR technology.
[0174] (3) Business card OCR recognition;
[0175] (4) Document OCR recognition;
[0176] Based on scanning technology, paper documents such as books and newspapers can be digitized, and the current English recognition rate is also very high. In recent years, it has also been used for mobile document recognition, and can be recognized by scanning.
[0177] (5) Bill OCR recognition;
[0178] Bill OCR recognition is used for various bill recognition, based on template mechanism, different recognition elements need to be customized for different bills, this technology is also called element recognition OCR, the earliest application is in the banking industry, many enterprises, financial and telecommunications institutions are using it.
[0179] (6) License plate OCR recognition;
[0180] Step S304, the first analysis result of the target business document and the second analysis result of the target business document are used as the first analysis result of the target business document.
[0181] Specifically, as can be known from the above introduction, the method provided by the embodiments of the present application can determine to use different processing methods to process the target business document by judging whether there is a document or picture that can be parsed in the target business document, so as to obtain the first analysis result and the second analysis result.
[0182] The first analysis result and the second analysis result include target information in the target business document.
[0183] Therefore, after obtaining the first analysis result and the second analysis result, the first analysis result of the target business document and the second analysis result of the target business document can be used as the first analysis result of the target business document.
[0184] As can be seen from the above introduction, the method provided by the embodiments of the present application can determine to use different processing methods to process the target business document by judging whether there is a document or picture that can be parsed in the target business document, so as to obtain the first analysis result and the second analysis result as the first analysis result of the target business document, so that the target information can be obtained by analyzing the first analysis result of the target business document.
[0185] As introduced above, the method provided by the embodiments of the present application can parse the documents or pictures that can be parsed in the target business document, and convert the documents that can be parsed into pictures to obtain the first analysis result of the target business document. Next, the process is introduced, which can include the following steps:
[0186] In step S401, the documents or pictures that can be parsed in the target business document are parsed to obtain the parsing result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document.
[0187] Specifically, as introduced above, the method provided by the embodiments of the present application can determine whether the target business document has documents or pictures that can be parsed and documents or pictures that cannot be parsed. When the target business document has documents or pictures that can be parsed, it is indicated that the documents or pictures that can be parsed in the target business document can be obtained by parsing processing.
[0188] Therefore, when the target business document has documents or pictures that can be parsed, the documents or pictures that can be parsed in the target business document can be parsed, and the parsing result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document can be obtained.
[0189] In step S402, the documents that can be parsed are converted into pictures to obtain the pictures corresponding to the target business document.
[0190] Specifically, as introduced above, the method provided by the embodiments of the present application can parse the documents or pictures that can be parsed in the target business document to obtain the parsing result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document.
[0191] After the target business document is parsed, the documents that can be parsed can be further converted into pictures to obtain the pictures corresponding to the target business document.
[0192] The pictures corresponding to the target business document include the image features corresponding to the target business document.
[0193] So that the picture corresponding to the target business document can be input into the preset deep instance segmentation model for batch processing to obtain target information included in the document or picture that can be parsed existing in the target business document.
[0194] Step S403, taking the image features of the picture corresponding to the target business document, the parsing result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document as the first analysis result of the target business document.
[0195] Specifically, as known from the above description, the method provided in the embodiments of the present application can parse the document or picture that can be parsed in the target business document after determining that the document or picture that can be parsed exists in the target business document.
[0196] After parsing the document or picture that can be parsed in the target business document to obtain the parsing result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document, and converting the document that can be parsed into a picture to obtain the picture corresponding to the target business document, the image features of the picture corresponding to the target business document, the parsing result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document can be further taken as the first analysis result of the target business document.
[0197] As can be seen from the above technical solutions, the method provided in the embodiments of the present application can parse the document or picture that can be parsed in the target business document, and convert the document that can be parsed into a picture to obtain the first analysis result of the target business document, so that the target information included in the target business document can be acquired by analyzing the first analysis result.
[0198] As known from the above description, the method provided in the embodiments of the present application can perform content recognition on the document or picture that cannot be parsed in the target business document, and convert the document that cannot be parsed into a picture to obtain the second analysis result of the target business document. Next, the process is introduced, which can include the following steps:
[0199] Step S501, performing content recognition on the document or picture that cannot be parsed in the target business document to obtain the recognition result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document.
[0200] Specifically, as introduced above, the method provided in the embodiments of the present application can determine whether the target business document exists a document or picture that can be parsed or a document or picture that cannot be parsed. When the target business document exists a document or picture that cannot be parsed, it indicates that the document or picture that cannot be parsed in the target business document cannot be processed by parsing to obtain the target information included in the document or picture that cannot be parsed in the target business document. The document or picture that cannot be parsed in the target business document needs to be processed by recognition to obtain the target information included in the document or picture that cannot be parsed in the target business document.
[0201] Therefore, when the target business document exists a document or picture that cannot be parsed, the content of the document or picture that cannot be parsed in the target business document can be recognized to obtain the recognition result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document.
[0202] In step S502, the document that cannot be parsed is converted into a picture to obtain the picture corresponding to the target business document.
[0203] Specifically, as introduced above, the method provided in the embodiments of the present application can recognize the document or picture that cannot be parsed in the target business document to obtain the parsing result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document.
[0204] After the target business document is processed by recognition, the document that cannot be parsed can be further converted into a picture to obtain the picture corresponding to the target business document.
[0205] The picture corresponding to the target business document includes the image features of the target business document.
[0206] So that the picture corresponding to the target business document can be input into the preset deep instance segmentation model for batch processing to obtain the target information included in the document or picture that cannot be parsed in the target business document.
[0207] In step S503, the image features of the picture corresponding to the target business document, the recognition result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document are taken as the second analysis result of the target business document.
[0208] Specifically, as introduced above, the method provided by the embodiments of the present application can perform recognition processing on the documents or pictures in the target business document that can be parsed after it is determined that the target business document contains documents or pictures that cannot be parsed.
[0209] After the documents or pictures in the target business document that cannot be parsed are parsed to obtain the parsing result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document, and the documents that cannot be parsed are converted into pictures to obtain the pictures corresponding to the target business document, the image features of the pictures corresponding to the target business document, the parsing result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document can be further taken as the second analysis result of the target business document.
[0210] As can be seen from the technical solutions introduced above, the method provided by the embodiments of the present application can perform recognition processing on the documents or pictures in the target business document that cannot be parsed, and convert the documents that cannot be parsed into pictures to obtain the second analysis result of the target business document, so that the target information included in the target business document can be obtained by analyzing the second analysis result.
[0211] In actual application, the purpose of parsing a business document can be to obtain the name of each business field corresponding to the target business document.
[0212] As introduced above, the method provided by the embodiments of the present application can input the first target information of the target business document into a preset form language processing model for processing to obtain the second target information of the target business document. Next, the process is introduced, which can include the following steps:
[0213] Step S701: Extract the name of each business field corresponding to the target business document in the first target information of the target business document to obtain the name of each business field corresponding to the target business document.
[0214] Specifically, as introduced above, the method provided by the embodiments of the present application can obtain the first target information included in the target business document by performing parsing processing on the target business document.
[0215] The first target information includes all related information of the target business document.
[0216] In order to further screen the target information, the first target information corresponding to the target business document can be further analyzed, so that the target information corresponding to the target business document can be quickly obtained.
[0217] Therefore, after obtaining the first target information, the name of each business field corresponding to the target business document in the first target information of the target business document can be further extracted, so as to obtain the name of each business field corresponding to the target business document.
[0218] So that the target information can be quickly obtained according to the name of each business field corresponding to the target business document.
[0219] In step S702, the name of each business field corresponding to the target business document is input into a preset form language processing model for business standard field alignment processing, so as to obtain the value corresponding to each business field in the target business document.
[0220] Specifically, as introduced above, the method provided by the embodiments of the present application can further extract the name of each business field corresponding to the target business document in the first target information of the target business document after obtaining the first target information, so as to obtain the name of each business field corresponding to the target business document.
[0221] As introduced above, the method provided by the embodiments of the present application can create a form language processing model to process the information related to the target business document.
[0222] Therefore, after obtaining the name of each business field corresponding to the target business document, the name of each business field corresponding to the target business document can be further input into a preset form language processing model for business standard field alignment processing, so as to obtain the value corresponding to each business field in the target business document.
[0223] For example, the text information included in the border of the region block corresponding to the key in a certain business document can be input into a preset form language processing model for processing.
[0224] For example, in an actual business form, there is text information such as "billing subject: ABCD Co., Ltd.", which can be divided into two region blocks corresponding to the text after processing by using a preset deep instance segmentation model, which are the border of the region block corresponding to the text "billing subject" and the border of the region block corresponding to the text "ABCD Co., Ltd.".
[0225] And the border of the region block corresponding to the "billing subject" text is the border of the region block corresponding to the key text, and "ABCD Co. Ltd." is judged as the border of the region block corresponding to the value text, and the structure information of the two is the left and right adjacent relationship. Further, the text information of the border of the region block corresponding to the key text can be classified, input "billing subject", and output as the standard field name: "company".
[0226] From the above technical solutions, the method provided by the embodiments of the present application can input the first target information of the target business document into a preset form language processing model for processing, quickly obtain the second target information of the target business document from the first target information. Accelerate the analysis speed of the related information included in the target business document, so that the desired target information in the target business document can be quickly obtained.
[0227] The form processing device provided by the embodiments of the present application is described below. The form processing device described below can be correspondingly referred to the form processing method described above.
[0228] Referring to Figure 2 , Figure 2 A form processing device structure schematic diagram disclosed by the embodiments of the present application.
[0229] As Figure 2 shown, the form processing device can include:
[0230] The data acquisition unit 101 is configured to acquire a business document to be processed.
[0231] The data preprocessing unit 102 is configured to preprocess the business document to be processed to obtain a target business document.
[0232] The data analysis unit 103 is configured to analyze the target business document to obtain a first analysis result of the target business document.
[0233] The first data processing unit 104 is configured to input the first analysis result of the target business document into a preset deep instance segmentation model for processing to obtain a second analysis result of the target business document.
[0234] The second data processing unit 105 is configured to analyze the second analysis result of the target business document to obtain first target information of the target business document.
[0235] The third data processing unit 106 is configured to input the first target information of the target business document into a preset form language processing model for processing to obtain second target information of the target business document.
[0236] From the above technical solutions can be seen, when the need to process business documents, the device provided by the embodiment of the application can obtain the to-be-processed business document by using the data acquisition unit 101; so that the to-be-processed business document can be preprocessed by using the data preprocessing unit 102, so that the format of the to-be-processed business document can be converted into a preset format, so that the target business document can be obtained, after obtaining the target business document, the target business document can be parsed by using the data parsing unit 103, and the first parsing result of the target business document is obtained, the first parsing result of the target business document includes all information of the target business document; by analyzing the first parsing result of the target business document, all information of the target business document can be understood, therefore, after obtaining the first parsing result of the target business document, the first parsing result of the target business document can be input into a preset deep instance segmentation model by using the first data processing unit 104 for processing, and the second parsing result of the target business document is obtained, wherein, by using the preset deep instance segmentation model to analyze the first parsing result of the target business document, all information of the target business document can be understood; after obtaining the second parsing result of the target business document, the second parsing result of the target business document can be further analyzed by using the second data processing unit 105, and the first target information of the target business document is obtained; so that the first target information of the target business document can be input into a preset form language processing model by using the second data processing unit 106 for processing, and the second target information of the target business document is obtained.
[0237] The device provided by the embodiment of the application can preprocess business documents of different formats, and analyze some important note information and unexpected situation description information in the business document, thereby all information of the business document can be analyzed, which helps to extract target information from all information of the business document, and the language processing model is used for further processing of the target information, not only the form information can be obtained by quickly processing the form, but also the obtained target information meets the preset format requirement, and can be applied to form processing in different scenes.
[0238] Further optionally, the data preprocessing unit 102 can include:
[0239] The cleaning unit is configured to perform data cleaning on the to-be-processed business document, and eliminate the to-be-processed business document with incorrect or incomplete content, to obtain a second to-be-processed business document;
[0240] The format identification unit is configured to identify the format of the second to-be-processed business document;
[0241] A format conversion unit is configured to convert the format of the second to-be-processed business document into a format meeting preset format requirements, so as to obtain the target business document.
[0242] Further, the data analysis unit 103 can include:
[0243] A judging unit is configured to judge whether the target business document contains analyzable documents or pictures and unanalyzable documents or pictures.
[0244] A first analysis unit is configured to, when the execution result of the judging unit is that the target business document contains analyzable documents or pictures, analyze the analyzable documents or pictures in the target business document, convert the analyzable documents into pictures, and obtain a first analysis result of the target business document.
[0245] A second analysis unit is configured to, when the execution result of the judging unit is that the target business document contains unanalyzable documents or pictures, identify the content of the unanalyzable documents or pictures in the target business document, convert the unanalyzable documents into pictures, and obtain a second analysis result of the target business document.
[0246] A data analysis subunit is configured to take the first analysis result of the target business document and the second analysis result of the target business document as a first analysis result of the target business document.
[0247] Further, the first analysis unit can include:
[0248] A first processing unit is configured to analyze the analyzable documents or pictures in the target business document, so as to obtain a parsing result of each character of the target business document, coordinate information of each character of the target business document in the target business document, and a text display attribute of each character of the target business document.
[0249] A second processing unit is configured to convert the analyzable documents into pictures, so as to obtain a picture corresponding to the target business document.
[0250] A third processing unit is configured to take an image feature of the picture corresponding to the target business document, the parsing result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document as the first analysis result of the target business document.
[0251] Further, the second analysis unit can include:
[0252] a fourth processing unit, configured to perform content recognition on a document or a picture that cannot be parsed in the target business document, to obtain a recognition result of each character of the target business document, coordinate information of each character of the target business document in the target business document, and a text display attribute of each character of the target business document;
[0253] a fifth processing unit, configured to convert the document that cannot be parsed into a picture, to obtain a picture corresponding to the target business document;
[0254] a sixth processing unit, configured to take the image feature of the picture corresponding to the target business document, the recognition result of each character of the target business document, the coordinate information of each character of the target business document in the target business document, and the text display attribute of each character of the target business document as a second analysis result of the target business document.
[0255] Further optionally, the creating process of the preset deep instance segmentation model can include:
[0256] taking the parsing result corresponding to the training business document as a training sample, taking target information contained in the parsing result corresponding to the training business document as a sample label, and training to obtain the preset deep instance segmentation model.
[0257] The preset deep instance segmentation model is trained by taking the parsing result corresponding to the training business document as a training sample and taking target information contained in the parsing result corresponding to the training business document as a sample label.
[0258] The parsing result corresponding to the training business document includes an image feature of a picture corresponding to the training business document, a parsing result and a recognition result of each character of the training business document, coordinate information of each character of the training business document in the training business document, and a text display attribute of each character of the training business document.
[0259] The target information contained in the parsing result corresponding to the training business document includes a character classification of each character of the training business document and each business field name corresponding to the training business document, a value corresponding to each business field of the training business document, a content of descriptive text corresponding to the training business document, and a border of each business field corresponding to the training business document.
[0260] Further optionally, the first target information includes a name of each business field corresponding to the target business document, and the process of inputting the first target information of the target business document into the preset form language processing model to obtain the second target information of the target business document can include:
[0261] extracting the name of each business field corresponding to the target business document in the first target information of the target business document to obtain the name of each business field corresponding to the target business document.
[0262] The name of each business field corresponding to the target business document is input into a preset form language processing model for business standard field alignment processing, to obtain the value corresponding to each business field in the target business document.
[0263] Further, the creation process of the preset form language processing model can include:
[0264] The first target information corresponding to the training business document is taken as a training sample, and the second target information contained in the first target information corresponding to the training business document is taken as a sample label, to obtain the form language processing model through training.
[0265] The first target information corresponding to the training business document includes the analysis result and the recognition result of each character of the training business document corresponding to the training business document.
[0266] The second target information contained in the first target information corresponding to the training business document includes the value corresponding to each business field and the content of the descriptive text of the training business document corresponding to the analysis result and the recognition result of each character of the training business document corresponding to the training business document.
[0267] The specific processing procedure of each unit included in the form processing apparatus can refer to the related description in the form processing method part, which will not be repeated here.
[0268] The form processing apparatus provided in the embodiments of the present application can be applied to a form processing device, such as a terminal: a mobile phone, a computer, etc. Optionally, A hardware structure block diagram of the form processing device is shown, which can refer to
[0269] The hardware structure of the form processing device can include at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4. Figure 3 Figure 3 In the embodiments of the present application, the number of processors 1, communication interfaces 2, memories 3, and communication buses 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 complete communication with each other through the communication bus 4.
[0270] The processor 1 can be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application, etc.
[0271] The processor 1 can be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application, etc.
[0272] The memory 3 can comprise a high-speed RAM memory and possibly also a non-volatile memory, such as at least one disk memory;
[0273] The memory stores a program, and the processor can invoke the program stored in the memory, and the program is used to implement each processing flow in the terminal form processing scheme.
[0274] The application further provides a readable storage medium, which can store a program suitable for processor execution, and the program is used to implement each processing flow in the terminal form processing scheme.
[0275] Finally, it should be noted that in this document, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or equipment including the element.
[0276] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the difference from other embodiments, and the same or similar parts between various embodiments can be referred to each other.
[0277] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the application. The various embodiments can be combined with each other. Therefore, the application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A form processing method, characterized in that, include: Retrieve pending business documents; The business documents to be processed are preprocessed to obtain the target business documents; The target business document is parsed to obtain the first parsing result of the target business document; The first parsing result of the target business document is input into a preset deep instance segmentation model for processing to obtain the second parsing result of the target business document; Analyze the second parsing result of the target business document to obtain the first target information of the target business document; The first target information of the target business document is input into a preset form language processing model for processing to obtain the second target information of the target business document. The preprocessing of the business document to be processed to obtain the target business document includes: Data cleaning is performed on the business documents to be processed to remove those with incorrect or incomplete content, resulting in a second business document to be processed. Identify the format of the second business document to be processed; The format of the second business document to be processed is converted into a format that meets the preset format requirements according to the preset format requirements, so as to obtain the target business document.
2. The method according to claim 1, characterized in that, The step of parsing the target business document to obtain the first parsing result of the target business document includes: Determine whether the target business document contains parseable documents or images, and whether it contains unparseable documents or images; If the target business document contains parsable documents or images, then the parsable documents or images in the target business document are parsed, and the parsable documents are converted into images to obtain the first analysis result of the target business document. If the target business document contains unparseable documents or images, then the content of the unparseable documents or images in the target business document is identified, and the unparseable documents are converted into images to obtain the second analysis result of the target business document. The first analysis result and the second analysis result of the target business document are used as the first parsing result of the target business document.
3. The method according to claim 2, characterized in that, The process of parsing the parsable documents or images in the target business document and converting the parsable documents into images to obtain the first analysis result of the target business document includes: The document or image in the target business document is parsed to obtain the parsing result of each character in the target business document, the coordinate information of each character in the target business document, and the text display attribute of each character in the target business document. Convert the parsable document into an image to obtain the image corresponding to the target business document; The image features of the image corresponding to the target business document, the parsing results of each character in the target business document, the coordinate information of each character in the target business document, and the text display attributes of each character in the target business document are used as the first analysis result of the target business document.
4. The method according to claim 2, characterized in that, The step of performing content recognition on unparseable documents or images in the target business document and converting the unparseable documents into images to obtain the second analysis result of the target business document includes: Content recognition is performed on documents or images in the target business document that cannot be parsed, to obtain the recognition result of each character in the target business document, the coordinate information of each character in the target business document, and the text display attributes of each character in the target business document; Convert unparseable documents into images to obtain the image corresponding to the target business document; The image features of the image corresponding to the target business document, the recognition result of each character in the target business document, the coordinate information of each character in the target business document, and the text display attributes of each character in the target business document are used as the second analysis result of the target business document.
5. The method according to claim 1, characterized in that, The creation process of the preset deep instance segmentation model includes: The training sample is obtained by using the parsing results corresponding to the training business document as training samples and the target information contained in the parsing results corresponding to the training business document as sample labels. in, The parsing results corresponding to the training business document include: the image features of the image corresponding to the training business document, the parsing and recognition results of each character of the training business document, the coordinate information of each character of the training business document in the training business document, and the text display attributes of each character of the training business document; The target information contained in the parsing results corresponding to the training business document includes: the character classification of each character in the training business document, the name of each business field corresponding to the training business document, the value of each business field corresponding to the training business document, the content of the descriptive text corresponding to the training business document, and the border of each business field corresponding to the training business document.
6. The method according to claim 1, characterized in that, The first target information includes the name of each business field corresponding to the target business document. The step of inputting the first target information of the target business document into a preset form language processing model for processing to obtain the second target information of the target business document includes: Extract the name of each business field corresponding to the target business document from the first target information of the target business document to obtain the name of each business field corresponding to the target business document. The name of each business field corresponding to the target business document is input into a preset form language processing model for business standard field alignment processing to obtain the value corresponding to each business field in the target business document.
7. The method according to claim 1, characterized in that, The creation process of the preset form language processing model includes: The training sample is obtained by using the first target information corresponding to the training business document as the training sample and the second target information contained in the first target information corresponding to the training business document as the sample label. in, The first target information corresponding to the training business document includes: the parsing result and recognition result of each character of the training business document; The second target information contained in the first target information corresponding to the training business document includes: the parsing result and recognition result of each character of the training business document, the value of each business field corresponding to each character, and the content of the descriptive text corresponding to the training business document.
8. A form processing device, characterized in that, include: The data acquisition unit is used to acquire business documents to be processed. The data preprocessing unit is used to preprocess the business documents to be processed to obtain the target business documents; The data parsing unit is used to parse the target business document and obtain the first parsing result of the target business document; The first data processing unit is used to input the first parsing result of the target business document into a preset deep instance segmentation model for processing to obtain the second parsing result of the target business document. The second data processing unit is used to analyze the second parsing result of the target business document to obtain the first target information of the target business document. The third data processing unit is used to input the first target information of the target business document into a preset form language processing model for processing, and obtain the second target information of the target business document. The data preprocessing unit includes: The cleaning unit is used to clean the data of the business documents to be processed, remove business documents with errors or incomplete content, and obtain a second business document to be processed. A format recognition unit is used to recognize the format of the second business document to be processed; The format conversion unit is used to convert the format of the second business document to be processed into a format that meets the preset format requirements, so as to obtain the target business document.
9. A form processing device, characterized in that, include: One or more processors, and memory; The memory stores computer-readable instructions that, when executed by the one or more processors, implement the steps of the form processing method as described in any one of claims 1 to 7.
10. A readable storage medium, characterized in that: The readable storage medium stores computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of the form processing method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for identifying information in picture, equipment and storage medium
CN112434689A
Structured information extraction method and device, electronic equipment and storage medium
CN113568965A
Form identification method and device, electronic equipment and computer readable medium
CN114612921A