Information Recognition Method, Device, Equipment, and Storage Medium

By using information identification methods in the financial reimbursement system, the errors in the identification of text information in financial reimbursement in Chinese are solved, and the traditional financial reimbursement problems are low efficiency and low accuracy are achieved, and a more efficient and accurate financial reimbursement process is achieved.

CN114973221BActive Publication Date: 2025-06-17ALIBABA GROUP HOLDING LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110215064.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-25
Publication Date
2025-06-17
Estimated Expiration
2041-02-25

AI Technical Summary

Technical Problem

Traditional financial reimbursement methods are inefficient, rely on manual review and entry, and the accuracy of text information recognition results is not high, resulting in incorrect reimbursement results.

Method used

An information recognition method is provided, by displaying the image of the target object in the interface, determining the information structure and marking information contained in the image, and outputting these information structures and marking information, helping users quickly locate and correct errors in the recognition results.

Benefits of technology

It improves the efficiency and accuracy of financial reimbursement, reduces the time and error rate of manual review, and makes the financial reimbursement process more automated and reliable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114973221B_ABST
    Figure CN114973221B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides an information recognition method, device, equipment and storage medium. The method includes: displaying an image of a target object in a first display area of an interface; determining at least one information structure included in the image and respective marking information corresponding to the at least one information structure, each information structure corresponding to a different field in the target object; and outputting the at least one information structure and the respective marking information corresponding to the at least one information structure in a second display area of the interface. A user can quickly distinguish each information structure based on the marking information corresponding to each information structure, so as to review and correct each information structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an information recognition method, device, equipment and storage medium. Background Art

[0002] Traditional financial reimbursement methods mainly involve reimbursement personnel submitting various cards, certificates, and bills offline to financial personnel, who manually complete the review, entry, and settlement of reimbursement information, wasting a large amount of manpower and time and having low efficiency.

[0003] With the development of artificial intelligence and cloud computing technologies, intelligent office has entered more and more enterprises. As an important part of intelligent office, financial reimbursement can automatically identify relevant text information in cards, certificates, and bills based on Optical Character Recognition (OCR) technology, realizing the electronic entry of relevant text information.

[0004] Affected by many factors, the accuracy of text information recognition results is not very reliable. If inaccurate text information is entered, it will lead to incorrect reimbursement results. Summary of the Invention

[0005] Embodiments of the present invention provide an information recognition method, device, equipment and storage medium, which can help users quickly locate possible errors in text information recognition results.

[0006] In a first aspect, embodiments of the present invention provide an information recognition method, which includes:

[0007] Display an image of a target object in a first display area of an interface;

[0008] Determine at least one information structure included in the image and respective marking information corresponding to the at least one information structure, where each information structure corresponds to a different field in the target object;

[0009] Output the at least one information structure and the respective marking information corresponding to the at least one information structure in a second display area of the interface.

[0010] In a second aspect, embodiments of the present invention provide an information recognition device, which includes:

[0011] A display module, configured to display an image of a target object in a first display area of an interface;

[0012] A detection module, configured to determine at least one information structure included in the image and respective marking information corresponding to the at least one information structure, where each information structure corresponds to a different field in the target object;

[0013] The display module is further configured to output the at least one information structure and the respective marking information corresponding to the at least one information structure in a second display area of the interface.

[0014] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory and a processor; wherein, an executable code is stored on the memory, and when the executable code is executed by the processor, the processor can at least implement the information recognition method as described in the first aspect.

[0015] In a fourth aspect, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which an executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the information recognition method as described in the first aspect.

[0016] In a fifth aspect, an embodiment of the present invention provides an information recognition method, and the method includes:

[0017] Receiving a request from a user device to call a target service interface, where the request includes an image of a target object;

[0018] Using processing resources corresponding to the target service interface to perform the following steps:

[0019] Determining at least one information structure included in the image and the respective marking information corresponding to the at least one information structure, and each information structure corresponds to a different field in the target object;

[0020] Sending the at least one information structure to the user device, so that the user device displays the image in a first display area of the interface and outputs the at least one information structure and the respective marking information corresponding to the at least one information structure in a second display area.

[0021] In a sixth aspect, an embodiment of the present invention provides an information recognition method, and the method includes:

[0022] Obtaining a form image;

[0023] Determining at least one information structure included in the form image and the respective marking information corresponding to the at least one information structure, and each information structure corresponds to a different field in the form image;

[0024] Outputting the at least one information structure and the respective marking information corresponding to the at least one information structure, for the user to complete information entry processing of the form image according to the marking information and the at least one information structure.

[0025] Seventh aspect, an embodiment of the present invention provides an information recognition method, which includes:

[0026] Obtain an image containing commodity information;

[0027] Determine at least one information structure included in the image and the respective marking information corresponding to the at least one information structure, and each information structure corresponds to a different field in the image;

[0028] Output the at least one information structure and the respective marking information corresponding to the at least one information structure, so that the user can check the entered commodity information according to the marking information and the at least one information structure.

[0029] Eighth aspect, an embodiment of the present invention provides an information recognition method, which includes:

[0030] Obtain a medical record image;

[0031] Determine at least one information structure included in the medical record image and the respective marking information corresponding to the at least one information structure, and each information structure corresponds to a different field in the medical record image;

[0032] Output the at least one information structure and the respective marking information corresponding to the at least one information structure, so that the user can screen out the medical record images that meet the requirements according to the marking information and the at least one information structure.

[0033] Ninth aspect, an embodiment of the present invention provides an information recognition method, which includes:

[0034] Obtain a teaching image;

[0035] Determine at least one information structure included in the teaching image and the respective marking information corresponding to the at least one information structure, and each information structure corresponds to a different field in the teaching image;

[0036] Output the at least one information structure and the respective marking information corresponding to the at least one information structure, so that the user can screen out the teaching images that meet the requirements according to the marking information and the at least one information structure. Tenth aspect, an embodiment of the present invention provides an information recognition method, which includes:

[0037] Obtain a reimbursement image containing a card certificate and / or a bill;

[0038] Determine at least one information structure included in the reimbursement image and the respective marking information corresponding to the at least one information structure, and each information structure corresponds to a different field in the reimbursement image;

[0039] Output the at least one information structure and the respective marking information corresponding to the at least one information structure, so that the user can complete the reimbursement process based on the marking information and the at least one information structure.

[0040] In an embodiment of the present invention, when an image containing a target object (such as a certain card or bill) is received, the text information in the image is recognized to obtain at least one information structure contained in the target object and the respective marking information corresponding to the at least one information structure. Among them, each information structure corresponds to a different field in the target object. For example, each information structure includes field attributes, field positions, and field contents. Finally, the recognized at least one information structure and the marking information corresponding to each information structure are output. In this way, the user can quickly view each information structure based on the marking information corresponding to each information structure, so as to correct the misrecognized information structure to ensure the accuracy of the final recognition result. Description of the Drawings

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 A schematic diagram of an image and an information structure provided by an embodiment of the present invention;

[0043] Figure 2 A flowchart of an information recognition method provided by an embodiment of the present invention;

[0044] Figure 3 A schematic diagram of the display effect of the information structure provided by an embodiment of the present invention;

[0045] Figure 4 A flowchart of another information recognition method provided by an embodiment of the present invention;

[0046] Figure 5 A flowchart of the information structure recognition process provided by an embodiment of the present invention;

[0047] Figure 6 A schematic diagram of the tabular information provided by an embodiment of the present invention;

[0048] Figure 7 A schematic diagram of the information structure recognition process provided by an embodiment of the present invention;

[0049] Figure 8 A schematic diagram of the display effect of the information structure provided by an embodiment of the present invention;

[0050] Figure 9 Flow chart of another information recognition method provided by an embodiment of the present invention;

[0051] Figure 10 Schematic application diagram of an information recognition method provided by an embodiment of the present invention;

[0052] Figure 11 Flow chart of another information recognition method provided by an embodiment of the present invention;

[0053] Figure 12 Flow chart of another information recognition method provided by an embodiment of the present invention;

[0054] Figure 13 Flow chart of another information recognition method provided by an embodiment of the present invention;

[0055] Figure 14 Flow chart of another information recognition method provided by an embodiment of the present invention;

[0056] Figure 15 Flow chart of another information recognition method provided by an embodiment of the present invention;

[0057] Figure 16 Schematic structural diagram of an information recognition device provided by an embodiment of the present invention;

[0058] Figure 17 For Figure 16 Schematic structural diagram of an electronic device corresponding to the information recognition device shown in the illustrated embodiment. Detailed implementation manners

[0059] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0060] The terms used in the embodiments of the present invention are merely for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. "Plural" generally includes at least two.

[0061] Depending on the context, as used herein, the terms "if" and "when" may be interpreted as "when", "while", "in response to determining", or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" may be interpreted as "when determined", "in response to determining", "when detected (stated condition or event)", or "in response to detecting (stated condition or event)".

[0062] In addition, the step timings in the following method embodiments are only examples and not strictly limited.

[0063] The information recognition method provided by the embodiments of the present invention can be executed by an electronic device, which can be a terminal device such as a PC, a laptop, a smart phone, etc., or a server. The server can be a physical server including an independent host, or can also be a virtual server, or can also be a server or server cluster in the cloud.

[0064] The information recognition method provided by the embodiments of the present invention can be used to recognize text information of each object included in an image and output the recognized text information in a structured manner. In addition, since there may be some incorrect text information recognition results during the automatic text information recognition process, in order to enable users to more quickly and conveniently locate the potentially incorrect text information recognition results, while outputting the text information in a structured manner, the corresponding marking information of the text information can also be output at the same time, so that users can quickly focus on and review the accuracy of each recognition result according to the marking information corresponding to different recognition results.

[0065] Among them, the structured output of text information here means outputting groups of text information recognized from the input image in the form of key-value pairs, that is, outputting each recognized information structure.

[0066] Taking the input image to be recognized as an image of a target object as an example, the target object can be, for example, a card, a bill, a statement, a product flyer, etc. The target object actually includes multiple fields, and each information structure mentioned herein corresponds to a different field in the target object. During the process of information recognition of the image, the relevant information of each field needs to be extracted, and the relevant information of each field can be represented in the form of a structure. It can be seen that an information structure is used to represent a field included in the target object. In addition, through the information structure, image format data (the input is image data) can also be converted into text format data (the information structure is in text format), so as to facilitate subsequent operations such as editing and information entry by users on the text format data.

[0067] In practical applications, each information structure is composed of a corresponding set of field attributes, field positions, and field contents. Among them, the field attributes and field positions can be regarded as keys, and the field contents can be regarded as key values. Since an information structure is used to represent a field in a target object, in order to accurately represent a field, in the embodiments of the present invention, the above three types of information are used to represent a field. Suppose the target object includes field X, the field position is the pixel position of field X in the image, and the field position will be automatically output during the process of identifying the information structure of the image. The field content, as the name implies, is the actual text content filled at field X, and it is automatically obtained through text recognition processing during the process of identifying the information structure. Since it is assumed in the embodiments of the present invention that the target object has a fixed plate type, each field in this type of target object has a specific physical meaning, and this physical meaning can be represented by an automatic attribute. In other words, the field attribute is used to describe the physical meaning of the field content.

[0068] The information structure can be represented in the following format: [field position, field attribute, field content]. Through this information structure, it can be known that there is a field at a certain position in the image of the target object, what the content of this field is, and what the physical meaning of this field is.

[0069] For ease of understanding, in combination with Figure 1 to exemplify the composition of an image and the meaning of the information structure.

[0070] In Figure 1 , it is assumed that an image includes an ID card and a train ticket as shown in the figure, that is, a photo is taken of a user's ID card and train ticket together to obtain an image. The purpose of performing text information recognition on this image is to extract each information structure included in the ID card image area and each information structure included in the train ticket image area.

[0071] Based on Figure 1 the above assumption, the information structures identified in the ID card image area may include, but are not limited to:

[0072] [L1, Name, Zhang XX], [L2, Date of Birth, September 20, 1990], [L3, Address, a certain community in a certain district of a certain city].

[0073] Among them, "L1" represents the field position corresponding to the name field in the ID card image area, "Name" is the field attribute (or field category, field name), and "Zhang XX" represents the field content.

[0074] Similarly, "L2" and "L3" respectively represent the field positions corresponding to the date of birth field and the address field in the ID card image area.

[0075] Similarly, the information structure recognized in the train ticket image area may include, but is not limited to:

[0076] [L4, starting station, Guangzhou East Station], [L5, terminal station, Beijing South Station], [L6, fare, 500 yuan].

[0077] As Figure 1 shown, based on the recognition results of the above information structure, the corresponding character recognition results of the ID card image area and the train ticket image area can be respectively output on the interface. Specifically, as Figure 1 shown, for example, when displaying multiple information structures recognized from the ID card image area, the positional relationship of the multiple information structures can be determined according to the field positions included in the multiple information structures. Based on this positional relationship, the field attributes and field contents included in each of the multiple information structures are displayed on the interface. Among them, the positional relationship of the multiple information structures is: the positional relationship of one information structure relative to other information structures. For example, Figure 1 as shown, the information structure of "date of birth" is located somewhere below the information structure of "name".

[0078] In practical applications, after obtaining the above information structure, the field attributes and field contents included in each information structure can be recorded for subsequent use. For example, in the scenario of financial reimbursement, the information structures recognized from various cards and bills in the image can be entered into the reimbursement system to facilitate subsequent reimbursement processing by financial personnel.

[0079] Figure 1 As shown, it is the case where correct character information recognition results are obtained after the input image is subjected to character information recognition processing. However, in practical applications, it is impossible to ensure that the recognition results are completely correct and reliable. Therefore, the recognition results can be displayed for manual review and error correction to ensure that the final obtained character information is accurate.

[0080] The execution process of the information recognition method provided in this article will be exemplarily described below in conjunction with the following embodiments.

[0081] Figure 2 is a flowchart of an information recognition method provided by an embodiment of the present invention. As Figure 2 shown, the method includes the following steps:

[0082] 201. Display an image of the target object in the first display area of the interface.

[0083] 202. Determine at least one information structure included in the image and the marking information corresponding to each of the at least one information structure, and each information structure corresponds to a different field in the target object.

[0084] 203. Output at least one information structure and the respective marking information corresponding to at least one information structure within the second display area of the interface.

[0085] In an embodiment of the present invention, the input image may be an image including at least one object, and each object may include multiple fields. The above target object may be any one of the at least one object, and the target object includes multiple fields.

[0086] In practical applications, at least one object including the target object may be photographed together to obtain an image of the above target object. In different application scenarios, the target object is different. For example, in a reimbursement scenario, the target object may be one or several cards, certificates, or bills. For example, in a medical scenario, the target object may be an electronic medical record. For example, in an e-commerce scenario, the target object may be a product flyer.

[0087] In practical applications, in different application scenarios, users may have a need to extract the text information contained in the image of the target object. At this time, after obtaining the image of the target object, the user may adopt the method provided in the embodiment of the present invention to obtain the text information contained in the image, that is, at least one information structure, and can quickly complete the review of the extracted text information.

[0088] To facilitate the user's review of the recognition result, the initial input image before recognition and the recognition result may be simultaneously displayed on the interface of the user device, where the initial input image is the image of the target object, and the recognition result is at least one information structure and the respective marking information corresponding to at least one information structure obtained by recognition.

[0089] Among them, the recognition process of the information structure can be obtained by a pre-trained network model, and the specific recognition process will be described in detail below.

[0090] Optionally, the marking information corresponding to the information structure may be determined according to the confidence of the information structure obtained during the process of recognizing the information structure. Or, optionally, the marking information corresponding to the information structure may also be determined according to the field position and field attribute included in the information structure. Or, the marking information corresponding to different information structures may be randomly determined.

[0091] Among them, when determining according to the field position and field attribute included in the information structure, the core idea is that the marking information corresponding to different information structures is different. For example, fields with adjacent positions may be given different marking information to be visually distinguishable from adjacent fields. For another example, fields with different field attributes may be given different marking information to distinguish different fields.

[0092] For ease of understanding, in conjunction with Figure 3 the execution process of this embodiment will be described by way of example.

[0093] As Figure 3 shown, assume that the image of the target object is an image including a train ticket and a taxi ticket. At this time, the target object includes the train ticket and the taxi ticket. The initially input image, that is, the image including the train ticket and the taxi ticket, can be displayed in the first display area of the interface of the user device.

[0094] In Figure 3 , taking the identification of the information structure included in the train ticket as an example, assume that Figure 3 a plurality of information structures shown in the second display area are obtained. Then, according to the field positions recorded in these multiple information structures, while keeping the relative position relationship between different information structures unchanged, in the second display area, the field attributes and field contents recorded in each information structure are respectively displayed in the form of key-value pairs. At the same time, according to the marking information corresponding to each information structure, the marking information is added to the display area of the corresponding field attribute and / or field content. Figure 3 In

[0095] In practical applications, the manifestation forms of the marking information can include various forms such as colors, graphics, symbols, etc.

[0096] Based on the marking information corresponding to each information structure, the user can focus on and distinguish different information structures, and compare them with the respective fields in the input image, so as to know whether the recognition result of the information structure is correct and correct the misrecognized information structure.

[0097] As described above, the marking information corresponding to the information structure reflects the credibility of the information included in the information structure, enabling the user to focus on the review of the information structure with low credibility. For the information structure with high credibility, the user can choose to trust the recognition result of the foregoing model. In this way, when an edit operation triggered by the user according to the marking information corresponding to each information structure for the target information structure is received, the edit operation is executed to correct the target information structure. Among them, the target information structure is any one of the obtained multiple information structures. Generally, the user will trigger an edit operation for the information structure with low credibility. Therefore, the target information structure will generally be the information structure with low credibility. The above edit operation is often an operation to correct the field content and field attributes.

[0098] In another alternative embodiment, if an image includes many objects that need to be processed for text information recognition, and the number of fields to be recognized in each object is relatively large, then the number of finally obtained information structures is also large. If the user manually checks the correctness of each information structure one by one, it will obviously be time-consuming and laborious.

[0099] Therefore, in the information recognition solution provided in the embodiments of the present invention, it is possible to assist the user in quickly locating the information structures with relatively low credibility, so that the user can focus on the review of the field content corresponding to the information structure with relatively low credibility, and skip the review of the field content corresponding to the information structure with relatively high credibility, thereby improving efficiency. Therefore, while identifying at least one information structure included in the image, it is also necessary to obtain the confidence level of each information structure, so as to determine the marking information corresponding to each information structure based on this confidence level, and reflect the credibility of the corresponding information structure through this marking information. In this way, the user can focus on the text information recognition results with relatively low credibility and review and correct them.

[0100] Figure 4 The flowchart of another information recognition method provided by the embodiments of the present invention is as Figure 4 shown, and this method includes the following steps:

[0101] 401. Determine at least one information structure included in the target object and the confidence level corresponding to each of the at least one information structure in the image of the target object. Each information structure includes field attributes, field positions, and field contents.

[0102] Among them, the confidence level of the information structure is used to reflect the accuracy or credibility of the recognition result of this information structure. For example, assume that an information structure is [L1, starting station, Guangzhou East Station], then the confidence level corresponding to this information structure can reflect the probability that the image of the target object actually contains this information structure.

[0103] 402. Determine the marking information corresponding to each of the at least one information structure according to the confidence level corresponding to each of the at least one information structure.

[0104] 403. Output at least one information structure and the marking information corresponding to each of the at least one information structure.

[0105] In the embodiments of the present invention, the image of the target object may be an image including at least one object, the above-mentioned target object may be any one of the at least one object, and each object may include multiple fields.

[0106] For example, in the scenario of financial reimbursement, the image of the target object may include at least one card or certificate and / or bill. That is to say, the above at least one object may be at least one card or certificate and / or bill.

[0107] In different application scenarios, when multiple objects are included in the same image, the types of these multiple objects included in the same image are often not completely the same.

[0108] In practical applications, users can place the required cards, certificates, and bills (such as ID cards, train tickets, taxi tickets, airplane tickets, invoices) together for taking a photo to obtain an image, and provide the image to relevant staff (such as financial staff). The relevant staff inputs the image into the function module provided with the information recognition method provided in the embodiments of the present invention to obtain the output result of the function module. Simply put, the output result is each information structure body with marked information. Furthermore, the relevant staff locates the information structure body with relatively low credibility based on the marked information, and confirms the accuracy of the field content included in the information structure body. The above function module can be an application program (APP) on the user terminal locally, or a function module in an APP, or a service provided by the cloud. The user calls the corresponding call interface of the service, uploads the image to the cloud, and receives the recognition results feedback by the cloud: at least one information structure body and the confidence level corresponding to each of the at least one information structure body.

[0109] Next, the process of determining at least one information structure body included in the target object and the confidence level corresponding to each of the at least one information structure body will be specifically described. As Figure 5 shown, the following steps may be included:

[0110] 501. Identify the category corresponding to the target object and the position area of the target object in the image through an object detection model.

[0111] 502. Input the target image area cropped according to the position area into a text recognition model to recognize at least one set of text recognition results through the text recognition model, and the confidence levels of the field positions and the confidence levels of the field contents in each set of text recognition results. Each set of text recognition results includes a field position and a field content.

[0112] 503. Input the target image area and at least one set of text recognition results into a layout recognition model corresponding to the category of the target object to output at least one information structure body through the layout recognition model, and the confidence level of the corresponding relationship between the field attribute and the field content. Each information structure body includes a field attribute, a field position, and a field content.

[0113] It can be seen that the following three models are required in the embodiments of the present invention: an object detection model, a text recognition model, and a layout recognition model.

[0114] In practical applications, structurally, all three of these models can be implemented as neural network models, such as a Convolutional Neural Network (CNN) model; a Residual Network (ResNet) model, such as ResNet-18, DLA-34 model, and so on.

[0115] Generally speaking, the object detection model is used to detect the category and location of each object in its input image, the text recognition model (a model obtained based on OCR technology) is used to recognize the location and content of each field in its input image, and the layout recognition model is used to recognize the field location and field attributes in its input image.

[0116] Among them, the layout recognition model is used to recognize the layout information contained in its input image, and the layout information reflects the location of each field in the input image and its corresponding field attributes. For example, cards and bills such as train tickets, invoices, and taxi tickets all have relatively fixed layout information. The layout recognition model corresponding to a certain card or bill is used to learn the layout information of the card or bill, so that it has the ability to extract the layout information of the card or bill.

[0117] For example, taking Figure 6 the train ticket shown in Figure 6 as an example, the layout information corresponding to the train ticket can be the field attributes and field locations of multiple fields included in the train ticket. In Figure 6 the field locations are represented by black rectangular graphics, and the field attributes are represented by the attribute labels associated with each field location, including Figure 6 the departure station, destination station, train number, seat number, departure date, fare, etc. shown in

[0118] For ease of understanding, the execution process of the embodiment shown in Figure 7 will be described in combination with the image shown in Figure 5 In Figure 7 it is assumed that the input image includes a train ticket and a taxi ticket. At this time, the target objects in the previous text can be the train ticket and the taxi ticket respectively. After inputting the image into the object detection model, the object detection model recognizes the object categories and object locations included in the input image, and thus obtains the recognition result: the input image includes two types of objects, one is a train ticket, and the other is a taxi ticket. The position area corresponding to the train ticket in the input image isFigure 7 In the area Q1 shown in, the position area corresponding to the taxi ticket in the input image is Figure 7 the area Q2 shown in.

[0119] After that, the train ticket image area and the taxi ticket image area can be intercepted according to the area Q1 and the area Q2 respectively, and the text recognition model is called to perform text recognition processing on the train ticket image area and the taxi ticket image area respectively. In this way, the positions of each field included in the train ticket image area and the field content at each field position can be obtained, and the positions of each field included in the taxi ticket image area and the field content at each field position can be obtained. The recognition results are as Figure 7 shown in.

[0120] It should be noted that taking the train ticket image area as an example, during the process of the text recognition model performing text recognition on the train ticket image area, on the one hand, it will output the positions of each field included in the train ticket image area and the field content at each field position, and on the other hand, it will also output the confidence level of each field position and the confidence level of each field content. Among them, the confidence level of a certain field position reflects the probability that there is indeed text at that position, and the confidence level of the field content reflects the probability that the text existing at the field position is indeed the recognized field content.

[0121] After that, if there is already a layout recognition model M1 corresponding to an object such as a train ticket, the train ticket image area, the field positions and the field content recognized by the text recognition model from the train ticket image area are input into the layout recognition model M1. The layout recognition model M1 outputs each information structure included in the train ticket image area. The composition of each information structure is: [field position, field attribute, field content]. The recognition result for the train ticket image area is as Figure 7 shown in. In addition, the layout recognition model M1 will also output the confidence level of the corresponding relationship between the field attribute and the field content in each information structure. This confidence level reflects the probability that the field attribute and the field content in an information structure match. In fact, this confidence level can be considered as the confidence level corresponding to the output result of the layout recognition model M1, and this output result is the information structure.

[0122] Suppose that after the train ticket image region is input into the text recognition model, the text recognition model identifies the field position Li and its corresponding field content Si therefrom (here only one field is taken as an example, and in fact, the positions and contents of all fields will be identified). After the train ticket image region, the field position Li, and the field content Si are input into the layout recognition model M1, the field content Si will be transparently transmitted and output. Based on the learned layout information of the train ticket, the layout recognition model M1 identifies the field position Li' and its corresponding field attribute Ci in the train ticket image region, as well as the confidence level corresponding to the recognition result, which is assumed to be Pa. The layout recognition model M1 can also determine the confidence level that the two positions corresponding to the field position Li and the field position Li' are the same field based on the distance between them, which is assumed to be Pb. If it is found that the field position Li corresponds to the field position Li', that is, it can be considered that the two correspond to the same field. Then, an information structure composed of the field position Li', the field attribute Ci, and the field content Si can be obtained, that is, the corresponding relationship between the field attribute Ci and the field content Si can be obtained. The confidence level of this corresponding relationship can be determined according to Pa and Pb, such as their product.

[0123] In fact, as Figure 7 shown, a certain field position L1 output by the text recognition model actually corresponds to the same field as the field position L1' output by the layout recognition model M1, and the same is true for other field positions. It can be considered that the layout recognition model M1 outputs a more accurate field position L1' based on the assistance of the field position L1 output by the text recognition model. In fact, the field positions included in the train ticket image region output by the text recognition model and the field content corresponding to each field position can reflect the different field positions in the train ticket image region and the relative position relationship between different field positions. Combining the field content corresponding to each field position can further reflect the semantic relationship between different field positions. Inputting the above output result of the text recognition model into the layout recognition model M1 can enable the layout recognition model M1 to output the layout information recognized from the train ticket image region more accurately based on the field positions, the relative position relationship between different field positions, and the semantic information between the field contents at different field positions: where the field of a certain attribute is located and what its field content is, that is, each information structure.

[0124] Similarly, if a layout recognition model M2 corresponding to an object such as a taxi ticket has been trained, then the taxi ticket image region and the field positions and field contents recognized by the text recognition model from the taxi ticket image region are input into the layout recognition model M2, and the output result of the layout recognition model M2 is as Figure 7 shown.

[0125] It can be seen from this that each object category corresponds to a layout recognition model.

[0126] It can be understood that the object detection model is trained to have the ability to recognize multiple object categories. When the object detection model detects what categories of objects are included in the input image, if it is found that the category of a certain object X cannot be recognized, it means that the object detection model has not been trained for this category during the training stage. When the object detection model, the text recognition model, and the layout recognition model are used as a whole, when the object detection model cannot recognize the category of object X, it can be considered that the layout recognition model corresponding to the category of object X also does not exist, and a layout recognition model corresponding to the category of object X needs to be trained. The training process of the layout recognition model will be introduced in other subsequent embodiments.

[0127] In the embodiments of the present invention, the object detection model can adopt an existing detection model that can be used to recognize multiple objects included in an image. The text recognition model can also adopt an existing model that can perform text recognition.

[0128] After obtaining at least one information structure included in the target object in the image and the confidence level corresponding to each information structure in the above manner, the marking information of the at least one information structure can be determined according to the confidence level corresponding to each of the at least one information structure. As can be seen from the above description, the confidence level corresponding to an information structure can be determined by the confidence level of the field position in the information structure, the confidence level of the field content, and the confidence level of the correspondence between the field attribute and the field content.

[0129] It is worth noting that various machine learning models and neural network models, including the object detection model, the text recognition model, and the layout recognition model provided in the embodiments of the present invention, can output a certain result and the confidence level corresponding to this result during their working processes. However, in the embodiments of the present invention, in order to finally obtain an accurate information recognition result and improve the information review and entry efficiency, it is also necessary to determine the marking information based on the confidence levels output by certain specific models (i.e., the several confidence levels mentioned above).

[0130] Specifically, for any information structure i recognized from the image, the confidence level P of the information structure i can be determined according to at least one of the following confidence levels:

[0131] The confidence level P1 of the field position in the information structure i, the confidence level P2 of the field content in the information structure i, and the confidence level P3 of the correspondence between the field attribute and the field content in the information structure i.

[0132] That is to say, the confidence level P of the information structure i can be determined according to the above three confidence levels P1, P2, and P3.

[0133] Optionally, the confidence level P of the information structure i may be determined according to a preset confidence weight, where the confidence weight includes at least one of the following: a first weight a1 corresponding to the confidence level of the field position, a second weight a2 corresponding to the confidence level of the field content, and a third weight a3 corresponding to the confidence level of the correspondence between the field attribute and the field content.

[0134] Specifically, the confidence level P of the information structure i may be determined as: P = a1 * P1 + a2 * P2 + a3 * P3.

[0135] In practical applications, the above three weights are preset values corresponding to three confidence levels respectively. Optionally, the third weight a3 may be set to be greater than or equal to the second weight a2, and the second weight a2 may be set to be greater than or equal to the first weight a1. Alternatively, the third weight a3 may also be set to be greater than the second weight a2 and the third weight a3 may be greater than the first weight a1. The reason for such setting is that the later obtained confidence level (i.e., the confidence level closer to the final output) has a greater impact on the final output result.

[0136] The above example considers the three confidence levels P1, P2, and P3 for the confidence level of an information structure i at the same time. In fact, some of these three confidence levels may also be used.

[0137] After obtaining the confidence level of the information structure i, the marking information corresponding to the information structure i is determined according to this confidence level.

[0138] Generally speaking, the corresponding relationship between different confidence level value ranges and marking information may be preset in advance, and the marking information corresponding to the information structure i is determined according to the value range in which the confidence level of the information structure i falls.

[0139] For example, optionally, two thresholds may be preset in advance: a first preset threshold and a second preset threshold, and the second preset threshold is greater than the first preset threshold. Through these two thresholds, three confidence level value ranges are determined: a range less than the first preset threshold: (0, the first preset threshold), a range composed of the first preset threshold and the second preset threshold: [the first preset threshold, the second preset threshold], and a range greater than the second preset threshold: (the second preset range, 1).

[0140] Thus, optionally, for any of the above information structures i, if the confidence level of the information structure i is less than the first preset threshold, the marking information corresponding to the information structure i is determined to be the first marking information; if the confidence level of the information structure i is between the first preset threshold and the second preset threshold, the marking information corresponding to the information structure i is determined to be the second marking information; if the confidence level of the information structure i is greater than the third preset threshold, the marking information corresponding to the information structure i is determined to be the third marking information.

[0141] Optionally, the above first marking information, second marking information, and third marking information may be of different colors. For example, the first marking information is green, the second marking information is yellow, and the third marking information is red.

[0142] Of course, optionally, the presentation form of the marking information may also adopt other forms, not limited to colors, such as different graphics, symbols, and so on.

[0143] In the embodiments of the present invention, the marking information corresponding to an information structure is used to reflect the credibility of the information structure. Briefly speaking, it reflects whether the recognition results of the field attributes, field contents, and field positions included in the information structure are accurate.

[0144] The above describes the determination process of the marking information of the information structure i. Based on this determination process, the marking information corresponding to each of the multiple information structures identified from the target image region corresponding to the target object can be obtained. After obtaining the marking information corresponding to each of the multiple information structures, multiple information structures and the marking information corresponding to each of the multiple information structures are output. In other words, multiple information structures are output according to the marking information corresponding to each of the multiple information structures. Among them, the marking information will affect the visual features of the information structure. For example, if the marking information corresponding to a certain information structure is red, the display area of the information structure will be rendered with a red background or highlighted in red. In this way, the user can determine its credibility according to the color of each information structure. For example, the credibility of red is the lowest, and the credibility of green is the highest. For the information structure with low credibility, compare it with the original input image and verify the correctness of the field content and field attributes in the information structure, and correct it in time when it is incorrect.

[0145] For ease of understanding, for Figure 7 the train ticket shown in the figure, in combination with Figure 8 to exemplarily illustrate the display result of the information structure.

[0146] As Figure 8 shown, optionally, the initially input image, that is, the image including the train ticket and the taxi ticket, can be displayed in the first display area of the interface. Since the object detection model will obtain the position areas of each object (in this example, the train ticket and the taxi ticket) included in the input image when detecting the input image, therefore, a selection box ( Figure 8 the bold selection box shown in the figure) can also be displayed in the first display area, as well as the control for this selection box: the buttons with the words "previous" and "next" shown in the figure. This selection box is used to frame an object according to the position areas of each object output by the object detection model, and the control is used to control the movement of the selection box so that the selection box switches between different objects.

[0147] In the second display area of the interface, the text information recognition result of the currently framed object can be displayed: each information structure and the corresponding marking information for each information structure.

[0148] In Figure 8 , assume that the currently selected bounding box selects a train ticket, and assume that multiple information structures as shown in Figure 7 are recognized from the train ticket image area. Then, according to the field positions recorded in these multiple information structures, while keeping the relative position relationship between different information structures unchanged, in the second display area, the field attributes and field contents recorded in each information structure are respectively displayed in the form of key-value pairs. At the same time, according to the marking information corresponding to each information structure, the marking information is added to the display area of the corresponding field attribute and / or field content. Figure 8 In

[0149] As described above, the marking information corresponding to the information structure reflects the credibility of the information contained in the information structure, enabling the user to focus on the review of information structures with low credibility. For information structures with high credibility, the user can choose to trust the recognition result of the aforementioned model. In this way, when an edit operation triggered by the user according to the marking information corresponding to each information structure for the target information structure is received, the edit operation is executed to correct the target information structure. Among them, the target information structure is any one of the multiple obtained information structures. Generally, the user will trigger an edit operation for an information structure with low credibility. Therefore, the target information structure will generally be an information structure with low credibility. The above edit operation is often an operation to correct the field content and field attributes.

[0150] In summary, based on each information structure contained in the target object obtained by performing text information recognition processing on the target object in the image and the confidence of each information structure, marking information can be set for the corresponding information structure according to the high or low confidence of each information structure to reflect the credibility of the information structure. In this way, the user can quickly locate the information structures that are prone to recognition errors and review and correct them in a timely manner, improving the efficiency compared to manually checking all information structures one by one.

[0151] Figure 9 It is a flowchart of another information recognition method provided by an embodiment of the present invention. As Figure 9 shown, the method may include the following steps:

[0152] 901. The object detection model fails to recognize the category corresponding to the target object in the image.

[0153] 902. Obtain the first training sample corresponding to the target object.

[0154] 903. Determine the annotation information of the first training sample, and train a plate recognition model based on the first training sample and the annotation information. The first annotation information includes each information structure included in the first training sample, and each information structure includes a field attribute, a field position, and a field content.

[0155] 904. Obtain a second training sample according to the image.

[0156] 905. Determine the annotation information of the second training sample, and train an object detection model based on the second training sample and the annotation information. The second annotation information includes the category and position area of each object included in the second training sample.

[0157] In this embodiment, the image refers to an image including at least one object, and the target object can be any one of the at least one object. For example, the at least one object is at least one card and / or bill.

[0158] In practical applications, for the object detection model, assume that initially, the object detection model is trained to have the ability to recognize N types of objects, where N is greater than or equal to 1. During subsequent use, new requirements may arise, such as hoping that the object detection model can also recognize the ability to recognize another certain type of object.

[0159] For example, the object detection model initially has the ability to recognize identity cards and train tickets. Later, there is a need to recognize taxi tickets. At this time, the initial object detection model will not be able to recognize taxi tickets, so the object detection model needs to be optimized to make it have the ability to recognize taxi tickets and the like.

[0160] In addition, as mentioned above, when the object detection model cannot recognize the category of an object, it can be considered that there is also no corresponding plate recognition model for this category. Therefore, in order to perform text information recognition processing on objects of this category subsequently, it is also necessary to train a corresponding plate recognition model for this category.

[0161] Based on this, when an image is input into the object detection model, the object detection model recognizes the category and position area of each object included in the image. The following situations may be encountered during the recognition process: The object detection model can recognize the position area of a certain object, but cannot recognize the category of the object. At this time, it will trigger the training of the plate recognition model corresponding to the category of the object, and the optimization training of the object detection model.

[0162] Suppose an object that cannot be recognized by the object detection model in an image is called a target object. For the training of the plate recognition model, first, a number of training sample images corresponding to the category of the target object need to be obtained, which are called the first training samples. Then, for supervised training, the first training samples need to be labeled with supervision information, which is called the first annotation information. Since the plate recognition model is used to learn the plate information of the target object, and the plate information can be reflected by the field positions and field attributes of multiple fields included in the target object, the first annotation information can include the field positions, field attributes, and field contents of each field included in the first training samples.

[0163] Among them, the field position can be represented by the four vertex coordinates of the rectangular box enclosing the field. Multiple field attributes included in the target object can be encoded in advance, so that the field attributes can be represented by the encoding results.

[0164] In practical applications, the acquisition method of the first training samples corresponding to the category of the target object is not specifically limited. It can be collected by the user independently or generated based on the adversarial network model.

[0165] Then, the first training samples and their corresponding first annotation information are input into the plate recognition model to train the plate recognition model. Generally speaking, based on the field attributes and field contents corresponding to different fields annotated in the first annotation information, the plate recognition model can learn the semantic information between different fields. Combining the field positions of each field, it can learn the relative field position relationship between fields with different field attributes, and finally enable the plate recognition model to have the ability to recognize the field positions and field attributes of different fields of the target object.

[0166] For the optimization training of the object detection model, first, a number of training sample images containing the target category (referring to the category of the target object) need to be obtained, which are called the second training samples. Then, for supervised training, the second training samples need to be labeled with supervision information, which is called the second annotation information. Since the object detection model is used to recognize the categories and positions of each object included in the image, the second annotation information can include the position areas and categories of each object included in the second training samples.

[0167] If the object detection model fails to identify the category C of object 1 in the input image X, then the result of obtaining the second training sample is: obtain an image of an object including category C, however, when the object detection model is only optimized and trained for category C, the image does not include other object categories that the object detection model cannot identify. Optionally, assume that the image X includes objects of category A, category B, and category C, and category C is the category that the object detection model cannot currently identify, so an image including objects of category A, category B, and category C can be collected as the second training sample.

[0168] After completing the training of the object detection model and the plate recognition model, for the subsequent input image of an object including category C, based on the solution provided in the foregoing embodiments, the recognition and display processing of the information structure in the object can be completed.

[0169] As described above, the information recognition method provided by the present invention can be executed in the cloud. Several computing nodes can be deployed in the cloud, and each computing node has processing resources such as computing and storage. In the cloud, a certain service can be organized by multiple computing nodes. Of course, a single computing node can also provide one or more services. The way for the cloud to provide the service can be to provide a service interface externally, and the user calls the service interface to use the corresponding service. The service interface includes forms such as a Software Development Kit (SDK) and an Application Programming Interface (API).

[0170] For the solution provided in the embodiments of the present invention, the cloud can provide a service interface for information recognition service, which is called the target service interface. When the user needs to perform information recognition on a certain image, the user device calls the target service interface to trigger a request to call the target service interface to the cloud, and the image of the target object to be recognized is carried in the request. The cloud determines the computing node that responds to the request, and uses the processing resources in the computing node to perform the following steps:

[0171] Determine at least one information structure included in the image and the respective marking information corresponding to at least one information structure, and each information structure corresponds to a different field in the target object;

[0172] Send at least one information structure to the user device, so that the user device displays the image in the first display area of the interface and outputs at least one information structure and the respective marking information corresponding to at least one information structure in the second display area.

[0173] For the detailed process of the target service interface using processing resources to perform information recognition processing, reference can be made to the relevant descriptions in the foregoing other embodiments, which will not be elaborated here. Additionally, it can be understood that the object detection model, text recognition model, and layout recognition model mentioned in the foregoing embodiments can run on one or different computing nodes in the cloud.

[0174] For ease of understanding, in combination with Figure 10 for exemplary illustration. In Figure 10 , when the user wants to perform information recognition processing on an image of a target object, the target service interface is called in the user device E1 to send a call request to the cloud computing node E2, and the call request includes the image of the target object. In this embodiment, it is assumed that the object detection model, text recognition model, and layout recognition model are running in the cloud computing node E2. Through these network models, the cloud computing node E2 identifies at least one information structure included in the image and the respective marking information corresponding to each information structure, and feeds back the recognition result to the user device E1. The user device E1 displays the at least one information structure and the respective marking information corresponding to each information structure on the interface for the user to operate on the information structure based on the marking information. In practical applications, the problem of image information recognition may be involved in many application fields, and the technical solutions of the embodiments of the present invention can be used.

[0175] In the reimbursement scenario, to improve the efficiency of financial reimbursement and the persistent storage management of reimbursement information, and at the same time facilitate the reimbursement process of users, it is not necessarily required to carry reimbursement-related cards, certificates, and bills to the financial personnel to conduct reimbursement. The solution provided by the embodiments of the present invention can be used to assist in completing the reimbursement process.

[0176] Figure 11 As shown in Figure 11 , which is a flowchart of another information recognition method provided by the embodiments of the present invention, the method may include the following steps:

[0177] 1101. Obtain a reimbursement image containing cards, certificates, and / or bills.

[0178] 1102. Determine at least one information structure included in the reimbursement image and the respective marking information corresponding to each information structure, and each information structure corresponds to a different field in the target object.

[0179] 1103. Output at least one information structure and the respective marking information corresponding to each information structure for the user to complete the reimbursement process based on the marking information and the at least one information structure.

[0180] Among them, users who need to claim reimbursement can take pictures of the cards and vouchers required for reimbursement placed together to obtain a reimbursement image containing these cards and vouchers, and transmit the reimbursement image online to the financial personnel. The financial personnel call the service interface providing information recognition services to upload the reimbursement image. By respectively recognizing the information structure bodies contained in each card and voucher in the reimbursement image, the information structure bodies contained in each card and voucher can be obtained. Each information structure body can be composed of a field position, a field attribute, and a field content. At the same time, while recognizing the information structure body, determine the corresponding marking information for each information structure body, and display each information structure body with the corresponding marking information, so that the financial personnel can focus on each information structure body according to the marking information corresponding to each information structure body, check the accuracy of the recognition result of each information structure body, perform a correction operation on the information structure body with incorrect recognition among them, and store each corrected information structure body. At this time, since each information structure body is data in text format, the financial personnel can perform operations such as copying and editing on the field content contained in each information structure body as needed. In addition, after obtaining each corrected information structure body, the field attributes and field contents included in the information structure body can also be stored in the reimbursement database in the form of key-value pairs according to the set storage policy.

[0181] In practical applications, in some application scenarios, paper forms such as tax bills and customs declarations are used. These forms usually also have a fixed format. Paper forms are not conducive to long-term storage. Therefore, it is necessary to digitally convert the paper forms for long-term storage. And digital conversion does not mean simply taking pictures of the paper forms to obtain the corresponding images. Because in practical applications, there may be a need to count and analyze form data. To support these needs, it is necessary to store the data content contained in the paper forms in text form. At this time, the information recognition solution provided by the embodiments of the present invention can be used to achieve this.

[0182] Figure 12 The flowchart of another information recognition method provided by the embodiments of the present invention is as Figure 12 shown. This method may include the following steps:

[0183] 1201. Obtain a form image.

[0184] 1202. Determine at least one information structure body contained in the form image and the corresponding marking information for each of the at least one information structure body. Each information structure body corresponds to a different field in the form image.

[0185] 1203. Output at least one information structure body and the corresponding marking information for each of the at least one information structure body for the user to complete the information entry process of the form image according to the marking information and the at least one information structure body.

[0186] Taking a photo of a certain form can obtain a form image. A form contains multiple cells, and each cell can be regarded as a field. Thus, an information structure can correspond to a cell.

[0187] After identifying the information structure corresponding to each cell from the form image and determining the marking information corresponding to each information structure, each information structure can be displayed according to the marking information corresponding to each information structure. Each information structure includes the position, attributes, and content of each cell.

[0188] Based on the marking information of each information structure, the user distinguishes different information structures, refers to the form image to check each information structure, and corrects the information structures with errors. After that, the cell attributes and cell contents included in each corrected information structure can be stored in the database in the form of key-value pairs. Based on this, assuming that a certain form has the attribute of payment amount, based on the storage results of a large number of forms, the user can trigger a statistical operation such as searching for forms with a payment amount greater than a certain set amount in the database to obtain the number of forms that meet the condition that the payment amount is greater than the set amount.

[0189] In the field of e-commerce, there is also a need for information extraction. For example, when a merchant has a new product to release, the merchant will register product information on the e-commerce platform (i.e., enter product information into the product database), and may also create a product promotion image to promote the product on the product interface so that consumers can understand the product more comprehensively and in detail. Among them, the product promotion image includes product information. When the merchant registers product information, some product information may be entered incorrectly. At this time, by identifying the product information in the product promotion image, it can help the staff of the e-commerce platform check whether the product information entered by the merchant is incorrect.

[0190] Figure 13 The flowchart of another information recognition method provided by the embodiments of the present invention is as Figure 13 shown, and the method may include the following steps:

[0191] 1301. Obtain an image containing product information.

[0192] 1302. Determine at least one information structure included in the image and the marking information corresponding to each of the at least one information structure. Each information structure corresponds to a different field in the image.

[0193] 1303. Output at least one information structure and the marking information corresponding to each of the at least one information structure for the user to check the entered product information according to the marking information and the at least one information structure.

[0194] In this embodiment, the image containing product information can be the product promotion image in the above text. The product information may include multiple product attributes such as product name, model, size, color, price, manufacturer, etc. At this time, different product attributes serve as different field attributes, and the attribute value of each product attribute is the field content.

[0195] After identifying the information structure corresponding to each product attribute from the image and determining the marking information corresponding to each information structure, each information structure can be displayed according to the marking information corresponding to each information structure. Each information structure will include each product attribute, attribute value, and the position of each product attribute in the image.

[0196] The staff of the e-commerce platform distinguish different information structures based on the marking information of each information structure, query the corresponding product information in the database. The product information contains various product attributes and their attribute values that have been entered into the database. Compare the product information entered into the database with the product information revealed in the information structure to determine whether the product information entered into the database is correct, and correct the incorrect product attributes and attribute values.

[0197] In the medical field, a large number of electronic medical record images can be generated. When an institution needs to perform statistics and analysis on medical records, it can, with the assistance of the information recognition solution provided by the embodiments of the present invention, extract some key information contained in the medical record images, so as to perform in-depth management and analysis on the medical record images based on these key information.

[0198] Figure 14 It is a flowchart of another information recognition method provided by the embodiments of the present invention. As Figure 14 shown, the method may include the following steps:

[0199] 1401. Obtain a medical record image.

[0200] 1402. Determine at least one information structure contained in the medical record image and the marking information corresponding to each of the at least one information structure. Each information structure corresponds to a different field in the medical record image.

[0201] 1403. Output at least one information structure and the marking information corresponding to each of the at least one information structure for the user to screen out the medical record images that meet the requirements according to the marking information and the at least one information structure.

[0202] The medical record image can be an image obtained by photographing a fixed-plate medical record. The fields contained in the medical record image may include multiple fields corresponding to the basic information of the user and multiple fields related to diagnosis.

[0203] By recognizing the medical record images, the information structure corresponding to each field and the marking information corresponding to each information structure included therein can be obtained. Different information structures are distinguished based on the marking information of each information structure, and the misrecognized information structures are corrected with reference to the medical record images.

[0204] Information storage processing can be performed on each corrected information structure. For example, according to the field attributes and field contents included in each information structure, the information included in the medical record images is stored in a database in the form of key-value pairs, and this database can be a relational database.

[0205] After that, by performing data query processing on the database as needed, the statistical analysis requirements for various diseases can be realized. For example, taking the visit time and a certain disease as query keywords (the medical record images include fields corresponding to the visit time and fields corresponding to the disease type), the medical record images corresponding to this disease generated within a set period of time can be queried to view these medical record images.

[0206] In an educational scenario, a teacher may use demonstration tools such as blackboard writing and PPT during the teaching process, and students can take pictures of the demonstration tools to obtain teaching images. When students take a large number of teaching images, there is a need to classify and organize the large number of teaching images and retrieve them as needed later.

[0207] Figure 15 It is a flowchart of another information recognition method provided by an embodiment of the present invention. As Figure 15 shown, the method may include the following steps:

[0208] 1501. Obtain teaching images.

[0209] 1502. Determine at least one information structure included in the teaching image and the marking information corresponding to each of the at least one information structure, and each information structure corresponds to a different field in the teaching image.

[0210] 1503. Output at least one information structure and the marking information corresponding to each of the at least one information structure for the user to screen out the teaching images that meet the requirements according to the marking information and the at least one information structure.

[0211] A certain teacher may have his own writing style or PPT editing style, thus forming a relatively fixed layout feature. The teaching images may include various knowledge points presented in a tree-like relationship. For example, there are two parallel subheadings under a certain main heading. If the main heading is "Trigonometric Functions", the subheadings include "Sine", "Cosine", "Tangent", etc. In this embodiment, different fields in the teaching image may correspond to different knowledge point headings.

[0212] A information structure may include field attributes, field positions, and field contents. In this embodiment, the field attribute refers to the category corresponding to the knowledge point, the field content refers to the knowledge point name, and the field position refers to the pixel position corresponding to a certain knowledge point in the teaching image.

[0213] After obtaining at least one information structure and the respective marking information corresponding to the at least one information structure included in a certain teaching image through the foregoing information recognition scheme, the information structures can be displayed with the corresponding marking information, so that students can focus on what knowledge points are included in the current teaching image, and thus screen out the teaching images they need to view.

[0214] The above only takes several application fields as examples to illustrate the application scenarios applicable to the information recognition scheme provided by the embodiments of the present invention. In fact, it is not limited thereto.

[0215] The information recognition devices of one or more embodiments of the present invention will be described in detail below. Those skilled in the art can understand that these information recognition devices can all be configured by using commercially available hardware components through the steps taught by this solution.

[0216] Figure 16 As shown in the structure diagram of an information recognition device provided by an embodiment of the present invention, Figure 16 the device includes: a display module 11 and a detection module 12.

[0217] The display module 11 is used to display an image of a target object in a first display area of the interface.

[0218] The detection module 12 is used to determine at least one information structure included in the image and the respective marking information corresponding to the at least one information structure, and each information structure corresponds to a different field in the target object.

[0219] The display module 11 is further used to output the at least one information structure and the respective marking information corresponding to the at least one information structure in a second display area of the interface.

[0220] Optionally, the target object includes at least one card and / or bill.

[0221] Optionally, the device further includes an interaction module, which is used to execute a correction operation in response to a correction operation input by the user for a target information structure according to the respective marking information corresponding to the at least one information structure.

[0222] Optionally, the detection module 12 can specifically be used to: determine the confidence level corresponding to each of the at least one information structure; and determine the marking information corresponding to each of the at least one information structure according to the confidence level corresponding to each of the at least one information structure.

[0223] Each information structure includes a field attribute, a field position, and a field content. Thus, optionally, the detection module 12 can specifically be used to: for any one of the at least one information structure, determine the confidence level of the any one information structure according to at least one of the following confidence levels: the confidence level of the field position in the any one information structure, the confidence level of the field content in the any one information structure, the confidence level of the correspondence relationship between the field attribute and the field content in the any one information structure.

[0224] Optionally, the detection module 12 can specifically be used to: determine the confidence level of the any one information structure according to a preset confidence weight; wherein, the confidence weight includes at least one of the following: a first weight corresponding to the confidence level of the field position, a second weight corresponding to the confidence level of the field content, a third weight corresponding to the confidence level of the correspondence relationship between the field attribute and the field content.

[0225] Optionally, the third weight is greater than or equal to the second weight, and the second weight is greater than or equal to the first weight.

[0226] Optionally, the detection module 12 can specifically be used to: for any one of the at least one information structure, if the confidence level of the any one information structure is less than a first preset threshold, determine that the marking information corresponding to the any one information structure is a first marking information; if the confidence level of the any one information structure is between the first preset threshold and a second preset threshold, determine that the marking information corresponding to the any one information structure is a second marking information; if the confidence level of the any one information structure is greater than the second preset threshold, determine that the marking information corresponding to the any one information structure is a third marking information.

[0227] Optionally, the first marking information, the second marking information, and the third marking information are of different colors.

[0228] Optionally, the detection module 12 may specifically be configured to: identify the category corresponding to the target object and the position area of the target object in the image through an object detection model; input the target image area cropped according to the position area into a text recognition model to recognize at least one set of text recognition results through the text recognition model, where each set of text recognition results includes a field position and a field content; input the target image area and the at least one set of text recognition results into a layout recognition model corresponding to the category to output at least one information structure through the layout recognition model, and each information structure includes a field attribute, a field position, and a field content.

[0229] Optionally, the confidence level of the field position and the confidence level of the field content are output by the text recognition model, and the confidence level of the correspondence relationship between the field attribute and the field content is output by the layout recognition model.

[0230] Optionally, the device further includes: a first training module, configured to obtain a first training sample corresponding to the target object if the category corresponding to the target object is not recognized in the image through the object detection model; determine first annotation information of the first training sample, where the first annotation information includes each information structure included in the first training sample, and each information structure includes a field attribute, a field position, and a field content; train the layout recognition model according to the first training sample and the first annotation information.

[0231] Optionally, the device further includes: a second training module, configured to obtain a second training sample according to the image if the category corresponding to the target object is not recognized in the image through the object detection model; determine second annotation information of the second training sample, where the second annotation information includes the category and position area of each object included in the second training sample; train the object detection model according to the second training sample and the second annotation information.

[0232] Figure 16 The device shown can execute the information recognition method provided in the foregoing embodiments. For the detailed execution process and technical effects, refer to the descriptions in the foregoing embodiments, which will not be elaborated here.

[0233] In a possible design, the structure of the foregoing Figure 16 shown information recognition device can be implemented as an electronic device. As Figure 17 shown, the electronic device may include: a processor 21, a memory 22, and a display screen 23. Among them, executable code is stored on the memory 22. When the executable code is executed by the processor 21, the processor 21 can at least implement the information recognition method provided in the foregoing embodiments.

[0234] Optionally, the electronic device may further include a communication interface 24 for communicating with other devices.

[0235] In addition, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the information recognition method provided in the foregoing embodiments.

[0236] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0237] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, it can also be implemented by a combination of hardware and software. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0238] The information recognition method provided by the embodiment of the present invention can be executed by a certain program / software, which can be provided by the network side. The electronic device mentioned in the foregoing embodiments can download the program / software to the local non-volatile storage medium, and when it needs to execute the foregoing information recognition method, read the program / software into the memory through the CPU, and then the CPU executes the program / software to implement the information recognition method provided in the foregoing embodiments. The execution process can refer to the illustration in the foregoing Figures 1 to 15 above.

[0239] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An information recognition method, characterized in that, Including: Display an image of a target object in a first display area of the interface; Determine at least one information structure included in the image and respective marking information corresponding to the at least one information structure, each information structure corresponding to a different field in the target object; Output the at least one information structure and the respective marking information corresponding to the at least one information structure in a second display area of the interface; The determining at least one information structure included in the image includes: identifying, by an object detection model, a category corresponding to the target object and a position area of the target object in the image; inputting a target image area cropped according to the position area into a text recognition model to recognize at least one set of text recognition results through the text recognition model, where each set of text recognition results includes a field position and a field content; inputting the target image area and the at least one set of text recognition results into a layout recognition model corresponding to the category to output the at least one information structure through the layout recognition model, and each information structure includes a field attribute, a field position, and a field content.

2. The method according to claim 1, characterized in that, The method further includes: In response to a correction operation input by the user for a target information structure according to the marking information corresponding to the at least one information structure, execute the correction operation.

3. The method according to claim 1, characterized in that, The determining the respective marking information corresponding to the at least one information structure includes: Determine the respective confidence levels of the at least one information structure; Determine the respective marking information corresponding to the at least one information structure according to the respective confidence levels of the at least one information structure.

4. The method according to claim 3, characterized in that, Each information structure includes a field attribute, a field position, and a field content. The determining the respective confidence levels of the at least one information structure includes: For any one of the at least one information structures, determine the confidence level of the any one information structure according to at least one of the following confidence levels: The confidence level of the field position in the any one information structure, the confidence level of the field content in the any one information structure, the confidence level of the correspondence relationship between the field attribute and the field content in the any one information structure.

5. The method according to claim 4, characterized in that, The determining the confidence level of the any one information structure includes: Determine the confidence level of the any one information structure according to a preset confidence weight; Wherein, the confidence weight includes at least one of the following: a first weight corresponding to the confidence level of the field position, a second weight corresponding to the confidence level of the field content, a third weight corresponding to the confidence level of the correspondence relationship between the field attribute and the field content.

6. The method according to claim 5, characterized in that, The third weight is greater than or equal to the second weight, and the second weight is greater than or equal to the first weight.

7. The method according to claim 3, characterized in that, The determining the respective marking information corresponding to the at least one information structure according to the respective confidence levels of the at least one information structure includes: For any one of the at least one information structures, if the confidence level of the any one information structure is less than a first preset threshold, determine the marking information corresponding to the any one information structure as a first marking information; If the confidence level of any of the information structures is between the first preset threshold and the second preset threshold, determine that the marking information corresponding to any of the information structures is the second marking information; If the confidence level of any of the information structures is greater than the second preset threshold, determine that the marking information corresponding to any of the information structures is the third marking information.

8. The method according to claim 1, characterized in that, The method further includes: If the category corresponding to the target object is not recognized in the image by the object detection model, obtain the first training sample corresponding to the target object; Determine the first annotation information of the first training sample, where the first annotation information includes each information structure included in the first training sample, and each information structure includes a field attribute, a field position, and a field content; Train the form recognition model according to the first training sample and the first annotation information.

9. The method according to claim 1, characterized in that, The target object is at least one card and / or bill.

10. An information recognition device, characterized in that, Includes: A display module, configured to display an image of a target object in a first display area of the interface; A detection module, configured to determine at least one information structure included in the image and the marking information corresponding to each of the at least one information structure, and each information structure corresponds to a different field in the target object; The display module is further configured to output the at least one information structure and the marking information corresponding to each of the at least one information structure in a second display area of the interface; The detection module is specifically configured to: recognize the category corresponding to the target object and the position area of the target object in the image through an object detection model; input the target image area intercepted according to the position area into a text recognition model to recognize at least one set of text recognition results through the text recognition model, and each set of text recognition results includes a field position and a field content; Input the target image area and the at least one set of text recognition results into a form recognition model corresponding to the category, so as to output the at least one information structure through the form recognition model, and each information structure includes a field attribute, a field position, and a field content.

11. An electronic device, characterized in that, Includes: A memory and a processor; wherein, an executable code is stored on the memory, and when the executable code is executed by the processor, the processor executes the information recognition method according to any one of claims 1 to 9.

12. A non-transitory machine-readable storage medium, characterized in that, An executable code is stored on the non-transitory machine-readable storage medium, and when the executable code is executed by a processor of an electronic device, the processor executes the information recognition method according to any one of claims 1 to 9.

13. An information recognition method, characterized in that, Includes: Receive a request from a user device to call a target service interface, where the request includes an image of a target object; Use the processing resources corresponding to the target service interface to perform the following steps: Determine at least one information structure included in the image and the marking information corresponding to each of the at least one information structure, and each information structure corresponds to a different field in the target object; Send the at least one information structure to the user device, so that the user device displays the image in a first display area of the interface and outputs the at least one information structure and the respective marking information corresponding to the at least one information structure in a second display area; The determining the at least one information structure included in the image includes: identifying, by an object detection model, a category corresponding to the target object and a position area of the target object in the image in the image; inputting a target image area cropped according to the position area into a text recognition model to recognize, by the text recognition model, at least one set of text recognition results, where each set of text recognition results includes a field position and a field content; inputting the target image area and the at least one set of text recognition results into a layout recognition model corresponding to the category to output, by the layout recognition model, the at least one information structure, and each information structure includes a field attribute, a field position, and a field content.

14. An information recognition method, characterized in that, including: Obtain a reimbursement image including a card and / or a bill; Determine at least one information structure included in the reimbursement image and the respective marking information corresponding to the at least one information structure, and each information structure corresponds to a different field in the reimbursement image; Output the at least one information structure and the respective marking information corresponding to the at least one information structure for the user to complete the reimbursement process according to the marking information and the at least one information structure; The determining the at least one information structure included in the reimbursement image includes: identifying, by an object detection model, a category corresponding to the card and / or the bill and a position area of the card and / or the bill in the reimbursement image in the reimbursement image; inputting a target image area cropped according to the position area into a text recognition model to recognize, by the text recognition model, at least one set of text recognition results, where each set of text recognition results includes a field position and a field content; Input the target image area and the at least one set of text recognition results into a layout recognition model corresponding to the category to output, by the layout recognition model, the at least one information structure, and each information structure includes a field attribute, a field position, and a field content.

Citation Information

Patent Citations

  • Train ticket identification method, system and device, and medium

    CN111340035A

  • Image recognition data error correction method and device, computer equipment and storage medium

    CN111582169A

  • Mobile supplementation, extraction, and analysis of health records

    US10395772B1