Image information extraction method and apparatus
By acquiring configuration information and calling the target base extractor to extract information from the image to be recognized, the problem of high cost and low efficiency in image information extraction in the existing technology is solved, and flexible and efficient image information extraction is achieved, enabling the extraction of information in different formats and types.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
- Filing Date
- 2023-05-23
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies suffer from high costs and low efficiency in image information extraction, and template matching and training models are not universally applicable. These technologies are not suitable for different scenarios, resulting in high costs, low efficiency, and poor versatility.
By obtaining configuration information, the target basic extractor is invoked to extract information from the image to be recognized and determine each of the target texts in the image to be recognized. The configuration information is used to characterize the invocation requirement for at least one target basic extractor, and each target basic extractor is used to extract target text of the corresponding information type.
It reduces the cost of image information extraction, improves the efficiency of image information extraction, and realizes the flexibility and scalability of image information extraction. It can extract information in different formats and types, thus reducing the cost of image information extraction and improving the efficiency of image information extraction.
Smart Images

Figure CN116665222B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information technology, and specifically to an image information extraction method and apparatus. Background Technology
[0002] Currently, Key Information Extraction (KIE) tasks have a wide range of applications. Specifically, KIE refers to the structured extraction of specific text from images. Through KIE, information in any format, such as form information, invoice information, or certificate information, can be extracted from images.
[0003] Existing technologies typically employ template matching or related models to achieve the task of extracting key information. However, summarizing templates and training models are not only costly but also time-consuming. Furthermore, the summarized templates or trained models are usually only applicable to the current scenario and cannot be generalized across different scenarios. This results in high cost and low efficiency for image information extraction using existing technologies. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide an image information extraction method and apparatus to reduce the cost of image information extraction and improve the efficiency of image information extraction.
[0005] In a first aspect, embodiments of the present invention provide an image information extraction method, the method comprising:
[0006] Acquire the image to be recognized;
[0007] Obtain configuration information, which is used to characterize the invocation requirement for at least one target basic extractor, and each target basic extractor is used to extract target text of the corresponding information type;
[0008] Each of the target base extractors is invoked to extract information from the image to be identified to determine each of the target texts in the image to be identified.
[0009] Secondly, embodiments of the present invention provide an image information extraction device, the device comprising:
[0010] An image acquisition unit is used to acquire the image to be recognized.
[0011] A configuration information acquisition unit is used to acquire configuration information, which is used to characterize the invocation requirement for at least one target basic extractor, and each target basic extractor is used to extract target text of the corresponding information type.
[0012] The extraction unit is used to invoke each of the target basic extractors to extract information from the image to be identified to determine each of the target texts in the image to be identified.
[0013] Thirdly, embodiments of the present invention provide a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the method as described in any one of the first aspects.
[0014] Fourthly, embodiments of the present invention provide an electronic device, the device comprising:
[0015] Memory is used to store one or more computer program instructions;
[0016] A processor, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of the first aspects.
[0017] The image information extraction method of this invention, after acquiring the image to be identified and configuration information, invokes a target basic extractor to extract information from the image to be identified based on the configuration information to determine each target text in the image. The configuration information represents the invocation requirement for at least one target basic extractor, and each target basic extractor is used to extract target text of a corresponding information type. This method can reduce the cost and improve the efficiency of image information extraction. Attached Figure Description
[0018] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which:
[0019] Figure 1 This is a schematic diagram of an application system for the image information extraction method according to an embodiment of the present invention;
[0020] Figure 2 This is a schematic diagram of the system framework of the image information extraction method according to an embodiment of the present invention;
[0021] Figure 3 This is a flowchart of the image information extraction method according to an embodiment of the present invention;
[0022] Figure 4 This is a flowchart of the target text extraction method according to an embodiment of the present invention;
[0023] Figure 5 This is a schematic diagram of the image to be identified according to an embodiment of the present invention;
[0024] Figure 6 This is a schematic diagram of the image to be identified according to an embodiment of the present invention;
[0025] Figure 7 This is a flowchart of the target text extraction method according to an embodiment of the present invention;
[0026] Figure 8 This is a schematic diagram of the image to be identified according to an embodiment of the present invention;
[0027] Figure 9 This is a schematic diagram of the image information extraction device according to an embodiment of the present invention;
[0028] Figure 10 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0029] The present invention is described below based on embodiments, but the invention is not limited to these embodiments. In the detailed description of the invention below, certain specific details are described in detail. Those skilled in the art will fully understand the invention even without these details. To avoid obscuring the essence of the invention, well-known methods, processes, flows, elements, and circuits are not described in detail.
[0030] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.
[0031] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application documents should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".
[0032] In the description of this invention, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this invention, unless otherwise stated, "a plurality of" means two or more.
[0033] It should be noted that, in order to protect user privacy, the user information involved in the various embodiments of the present invention is obtained with the user's authorization, and the obtained user information will only be applied to the methods in the various embodiments of the present invention.
[0034] Figure 1 This is a schematic diagram of an application system for the image information extraction method according to an embodiment of the present invention. For example... Figure 1 As shown, the application system includes at least one terminal 11 and a server 12.
[0035] In this embodiment, each terminal 11 is a user-owned terminal, which can be a common terminal such as a mobile phone, computer, or tablet computer. Furthermore, a target program can be installed on each terminal 11, which can be used to realize information interaction between the terminal 11 and the server 12. In this embodiment, when a user needs to extract target text from an image, they can activate the target program and upload the image to be recognized and the edited configuration information to the server 12 through the target program.
[0036] Optionally, the target program can be a standalone application (App) or a mini-program based on a platform program. Specifically, the mini-program is a program developed based on an application within the platform program and used to perform corresponding operations. Compared to a standalone application, the mini-program does not require additional download and installation; users can invoke the mini-program within the platform program. In an optional implementation, the target program can also be a web application that runs on a webpage, etc., which will not be elaborated further here.
[0037] The server 12 is a general-purpose computing device, data processing device, or storage device. Further, the server 12 may store at least one base extractor (BE), each of which can be used to extract target text of a corresponding information type. In this embodiment, the configuration information edited by the user can be used to characterize the call request for at least one target base extractor. After receiving the image to be recognized and the configuration information uploaded by the terminal 11, the server 12 can call at least one target base extractor according to the configuration information to extract the target text in the image to be recognized.
[0038] Specifically, developers can pre-define information extraction methods for different information types into corresponding basic extractors based on information type, and store each basic extractor in server 12. In practical applications, when a user needs to extract specific target text from an image to be recognized, they can convert their information extraction requirements for the image into a request to call at least one target basic extractor to edit and determine the corresponding configuration information. After editing the configuration information, the user can upload the image to be recognized and the edited configuration information to server 12 via terminal 11. After receiving the image to be recognized and the configuration information, server 12 can call at least one corresponding target basic extractor to extract the target text required by the user from the image to be recognized based on the configuration information.
[0039] It should be understood that each basic extractor may contain program code for implementing the corresponding information extraction method. Thus, the server can execute the corresponding information extraction method by calling the program code in each of the basic extractors.
[0040] Optionally, after extracting the target text, the server 12 may also send the target text back to the corresponding terminal 11 so that the terminal 11 can display the target text or perform other processing on the target text, which will not be elaborated here.
[0041] Optionally, the image information extraction method in this embodiment does not restrict the format of the extracted information. For example, the format can be a form, business card, ticket, or certificate.
[0042] Compared to existing technologies, this embodiment does not require additional summary templates or training models. Users can flexibly adjust the types of information they want to extract by editing configuration information, thereby reducing the cost of image information extraction and improving the efficiency of image information extraction.
[0043] Furthermore, the image information extraction method in this embodiment also has strong scalability. Specifically, this embodiment allows the image information extraction method to extract information of different formats and types by having developers update and extend the basic extractor in the server.
[0044] Optionally, the terminals 11 and the server 12 can be connected via a wireless network to achieve data interaction. The wireless network may include any one or a combination of 5G mobile communication network technology (5th-Generation, 5G), Long Term Evolution (LTE), Global System for Mobile Communication (GSM), Bluetooth (BT), Wireless Fidelity (Wi-Fi), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Range (Lora) wireless communication technology, or Zigbee protocol.
[0045] It should be understood that the image information extraction method in this embodiment is not limited to application on a server, but can also be applied to terminal devices similar to the terminal 11.
[0046] Figure 2 This is a schematic diagram of the system framework of the image information extraction method according to an embodiment of the present invention. Figure 2 As shown, the system framework includes a request processing layer 21, an image processing layer 22, a text recognition layer 23, a processing layer 24, an information extraction layer 25, and a target text output layer 26.
[0047] It should be understood that each layer in the system framework can be executed by the server device or terminal device in the above embodiments.
[0048] In the request processing layer 21, the device receives an information extraction request and obtains configuration information and the image to be identified based on the information extraction request.
[0049] In the image processing layer 22, the device performs image preprocessing on the image to be identified.
[0050] Optionally, the device may store at least one global image preprocessor (GP_IMG), each of which can be used to perform corresponding image preprocessing operations on the image to be recognized. Specifically, in this embodiment, the developers can pre-define the relevant image processing methods to form corresponding global image preprocessors and store each of the global image preprocessors in the device. This is illustrated by the image length scaling module, image cropping module, and image scaling module in image processing layer 22. Specifically, the image length scaling module is used to scale the image by a certain length, the image cropping module is used to crop the image, and the image scaling module is used to scale the image by a certain ratio.
[0051] Furthermore, users can characterize their invocation requirements for at least one target global image preprocessor by editing configuration information. In the image processing layer 22, the device will invoke the corresponding target global image preprocessor to perform image preprocessing on the image to be identified according to the configuration information. This can further improve the configurability of the image information extraction method, thereby better meeting the user's customized configuration requirements.
[0052] It should be understood that the image preprocessing operations may also include lighting and shadow processing operations, tilting operations, distortion operations, grayscale processing operations, or binarization operations, etc. The methods corresponding to each of the image preprocessing operations can be independently formed into corresponding global image preprocessors and stored in the device so that users can call them as needed, which will not be elaborated here.
[0053] Optionally, users can also define the relevant parameters involved in the execution of each of the aforementioned global image preprocessors through configuration information. For example, users can define scaling parameters through configuration information. It should be understood that when users do not define relevant parameters, each global image preprocessor can execute the corresponding method according to preset parameters.
[0054] In the text recognition layer 23, the device determines at least one text region to be recognized and the text content in each of the text regions to be recognized in the image to be recognized.
[0055] Optionally, the device may store at least one text processing module, each of which can be used to determine the text region to be recognized and the text content within each text region. Specifically, in this embodiment, the developer can pre-define methods such as the text region detection method, text direction classification method, and text content recognition method into separate text processing modules, and store each of these text processing modules separately in the device. This is illustrated by the text region detection module, text direction classification module, and text content recognition module in the text recognition layer 23. Specifically, the text region detection module is used to determine at least one text region to be recognized in the image to be recognized; the text direction classification module is used to detect the text direction within each text region to be recognized; and the text content recognition module is used to determine the text content within each text region to be recognized.
[0056] Furthermore, users can also characterize their need to invoke at least one text processing module by editing configuration information. In the text recognition layer 23, the device will invoke the corresponding text processing module according to the configuration information to determine the text region to be recognized and the text content in each of the text regions to be recognized.
[0057] In processing layer 24, if the configuration information includes filtering range information, the device will further filter the text region to be recognized based on the filtering range information. It should be understood that the filtering method can also be stored independently as a corresponding global preprocessor for OCR result (GP_OCR) on the server, so that the server can invoke it according to user needs.
[0058] Specifically, in some scenarios, multiple target texts of the same type may exist simultaneously in the same image to be recognized. In this case, in order to enable the user to obtain the target text they need, this embodiment allows the user to edit the filtering range information in the configuration information. In the processing layer 24, the device will filter the text region to be recognized within the defined filtering range according to the filtering range information to find the text region to be recognized that contains the target text needed by the user.
[0059] In the information extraction layer 25, the device will filter out the target text regions corresponding to each target basic extractor from multiple text regions to be identified, and use each target basic extractor to extract the target text in the corresponding target text region.
[0060] Optionally, each target basic extractor may have at least one corresponding keyword, and each keyword can be used to characterize the type of information in the target text to be extracted by the corresponding target basic extractor. When filtering target regions to be identified, the device can find the target text regions to be identified corresponding to each target basic extractor by determining whether the text content includes the corresponding keyword. For example, a name basic extractor used to extract name information may correspond to the keyword "name" or "Name". The device can determine that the text region to be identified includes the keyword "name" in its text content, and then identify that text region as the target text region to be identified corresponding to the name basic extractor.
[0061] Optionally, when the target text to be extracted by the target base extractor is a named entity, the target base extractor may also have a corresponding named entity type. Specifically, the named entity type can be the named entity type of the target text to be extracted by the target base extractor. When filtering target regions to be identified, the device can also find the target text regions to be identified corresponding to each target base extractor by determining whether the text content includes named entities of the corresponding type. Here, the named entity is specifically a proper noun used to refer to a person, place, or organization. For example, a name base extractor used to extract name information can correspond to a person's name. The device can determine that a text region to be identified includes a person's name in its text content, and then identify that text region as the target text region to be identified corresponding to the name base extractor.
[0062] It should be understood that the information extraction layer 25 can also combine the above two methods to find the text regions to be identified for each target, or the information extraction layer 25 can also use other methods, such as using a corresponding model to find the text regions to be identified for each target.
[0063] Optionally, when extracting target text, each of the target basic extractors can extract the target text based on the text features of the target text to be extracted. Specifically, each target text usually possesses corresponding text features. For example, for the target text of a telephone number, it is usually composed of 8 or 11 digits. Another example is the target text of specific address information, which is usually composed of structures such as province, city, county (district), town (street), and village. The target basic extractor can find the corresponding target text in the text content of the target text region by recognizing the text features.
[0064] Optionally, when extracting target text, if keywords exist in the target text region to be identified, each of the aforementioned target basic extractors can also extract the target text based on the positional relationship between the target text and the corresponding keywords. Specifically, the target text and the corresponding keywords usually have a corresponding positional relationship. For example, the target text and the corresponding keywords are usually horizontally adjacent, and they are usually separated by a semicolon or colon. By determining the positional relationship between the target text and the corresponding keywords, the target basic extractor can find the target text corresponding to the keywords in the text content of the target text region to be identified.
[0065] Optionally, when extracting target text, if the target text to be extracted is a named entity, the target base extractor can also directly extract the named entity of the corresponding type in the target area to be identified as the target text.
[0066] In one alternative implementation, each of the target base extractors can combine the above three methods to extract the target text, or each of the target base extractors can employ other methods, such as using a multimodal classification model to identify the target text regions. The multimodal classification model can be used to classify text to identify the target text to be extracted.
[0067] It should be understood that different target base extractors may use the same or different methods when extracting target text.
[0068] In the target text output layer 26, the device can output each of the target texts.
[0069] It should be understood that not all of the aforementioned levels and modules within each level will be executed by the device. Some levels and certain modules within each level will only be invoked and executed by the device after being called by the user through configuration information. For example, processing layer 24 will only be invoked and executed by the device if the configuration information includes filtering range information. Similarly, the image scaling module will only be invoked and executed by the device if the configuration information specifies a call request for the image scaling module.
[0070] In an alternative implementation, the server may also store user-customizable general extractors such as K2V (key-to-value) basic extractors, Context basic extractors, and Ner (named entity recognition) basic extractors. It should be understood that these definitions can be implemented by the user through editing configuration information.
[0071] Specifically, the K2V basic extractor allows users to define the keywords to be extracted. The K2V basic extractor will identify the text region containing the user-defined keywords as the target text region to be identified, and extract the target text corresponding to the keywords from the target text region to be identified.
[0072] The Context base extractor allows users to define the context content of the target text to be extracted. The Context base extractor will identify the text region to be recognized containing the context content defined by the user as the target text region to be recognized, and extract the target text between the context content in the target text region to be recognized.
[0073] The Ner basic extractor allows the user to define the type of the extracted named entities. The Ner basic extractor will identify the text region to be identified containing the named entities of the user-defined type as the target text region to be identified, and identify the named entity as the target text.
[0074] It should be understood that Figure 2 The system framework for the image information extraction method shown is for illustrative purposes only. In actual applications, the specific layers of the system framework and the modules executed in each layer can be set and adjusted by the developers according to actual needs.
[0075] Figure 3 This is a flowchart of an image information extraction method according to an embodiment of the present invention. Figure 3 As shown, the image information extraction method may specifically include the following steps:
[0076] It should be understood that the execution subject of the image information extraction method in this embodiment can be the server in the above embodiments, or it can be the terminal in the above embodiments.
[0077] S100: Obtain the image to be recognized.
[0078] Specifically, the device can acquire an image to be recognized. This image includes the target text to be extracted.
[0079] Optionally, the target text can be in any format, such as a form, business card, ticket, or certificate.
[0080] Optionally, when the device is a server, the image to be identified can be an image uploaded by a terminal. Specifically, the device can receive an information extraction request sent by the terminal. The information extraction request may include the image to be identified or a URL link of the image to be identified. The device can obtain the image to be identified from the information extraction request, or it can obtain the image to be identified by accessing the corresponding address based on the URL (uniform resource locator) link in the information extraction request.
[0081] Optionally, when the device is a terminal, the image to be identified can also be acquired by the terminal under the user's control through a corresponding image acquisition device, such as a camera.
[0082] S200, Obtain Configuration Information.
[0083] The configuration information is used to characterize the call requirement for at least one target basic extractor, and each target basic extractor is used to extract target text of the corresponding information type.
[0084] Specifically, when a user needs to extract specific target text from an image, they can convert their information extraction requirements for the image to be recognized into a call to the basic extractor to edit and determine the corresponding configuration information, which the device can then obtain.
[0085] Optionally, when the device is a server, the configuration information can be carried in the information extraction request, and the device can obtain the configuration information from the information extraction request.
[0086] Optionally, when the device is a terminal, the device can obtain the configuration information by manual input by the user.
[0087] Optionally, the configuration information can be edited by the user using any type of computer programming language such as C, C++, JAVA, or Python. In an optional implementation, the configuration information can also be determined by the user through settings on the corresponding page.
[0088] S300: Invoke each of the target base extractors to extract information from the image to be identified to determine each of the target texts in the image to be identified.
[0089] Specifically, after acquiring the image to be recognized and the configuration information, the device can call at least one corresponding target base extractor to extract information from the image to be recognized based on the configuration information to determine each of the target texts in the image to be recognized.
[0090] Optionally, the configuration information can also be used to characterize the invocation requirement for at least one target global image preprocessor. Before extracting information from the image to be identified, the device can also invoke at least one corresponding target global image preprocessor to perform image preprocessing on the image to be identified according to the configuration information.
[0091] Figure 4 This is a flowchart of the target text extraction method according to an embodiment of the present invention. Figure 4 As shown, the target text extraction method may specifically include the following steps:
[0092] It should be understood that the target text extraction method can be specifically used to implement the above step S300.
[0093] S310. Determine at least one text region to be identified and the text content in each of the text regions to be identified in the image to be identified.
[0094] Specifically, the device can determine at least one text region to be recognized in the image to be recognized, and determine the text content in each text region to be recognized. The text region to be recognized can refer to a region in the image to be recognized that contains text information.
[0095] Optionally, in this step, the device can determine the target pixels belonging to text in the image to be recognized. After determining each target pixel, the device can determine at least one set of target pixels by traversing each target pixel, and then determine the text region to be recognized corresponding to each set of target pixels.
[0096] Optionally, during the traversal, for each target pixel, the device can determine the horizontal and vertical distances between the target pixel and the nearest target pixel. When both the horizontal and vertical distances meet preset conditions, the target pixel and the nearest target pixel are assigned to the same target pixel set. Thus, the device can determine at least one target pixel set after the traversal is complete. Specifically, meeting the preset conditions may mean that the horizontal distance is less than or equal to a first preset distance threshold, and the vertical distance is less than or equal to a second preset distance threshold.
[0097] It should be understood that in practical applications, target text and its corresponding keywords usually appear together, and their positional relationship is typically horizontally adjacent. In this embodiment, by adjusting the first preset distance threshold and the second preset distance threshold, the device can determine the target text and its corresponding keywords to be identified within the same text region.
[0098] It should be understood that, in order to ensure that multiple target texts are not identified into the same text region to be recognized, the specific values of the second preset distance threshold and the first preset distance threshold can be set and adjusted by the developers according to actual needs, or they can be determined by the computer through continuous simulation calculations.
[0099] Optionally, when determining target pixels, the device can use a corresponding model to determine the probability that each pixel in the image to be recognized belongs to text, and determine the pixels with a probability higher than a preset probability threshold as target pixels.
[0100] Figure 5 This is a schematic diagram of the image to be identified according to an embodiment of the present invention. Figure 5 As shown, the image to be recognized 51 includes a health certificate 52. The health certificate 52 includes textual and image information such as the certificate title, name, gender, validity period, number, issuing authority, and age.
[0101] In step S310, the device can determine at least one text region to be recognized in the image 51 to be recognized, that is, the region enclosed by the dashed boxes. Each text region to be recognized may include a keyword and the target text corresponding to the keyword. For example, for the text region 53 to be recognized, the text region 53 to be recognized includes the keyword "name" and the specific name information corresponding to the keyword "name".
[0102] Optionally, after determining each of the text regions to be identified, the device can perform text recognition on each of the text regions to be identified to determine the text content in each of the text regions to be identified.
[0103] Furthermore, in order to successfully determine the text content, the device can also detect the text orientation of each of the text regions to be identified before performing text recognition on each of the text regions to be identified. After determining the text orientation of each text region to be identified, the device can rotate the text region to be identified according to the text orientation so that the text orientation of the text region to be identified is horizontal.
[0104] Optionally, before step S320, if the configuration information includes filtering range information, the device can also filter the text region to be recognized according to the filtering range information. Specifically, multiple target texts of the same information type may exist simultaneously in the same image to be recognized. In order to enable the user to obtain the target text they need, in this embodiment, the user can edit the filtering range information in the configuration information. The server will filter the text region to be recognized within the defined filtering range according to the filtering range information to find the text region to be recognized that contains the target text needed by the user.
[0105] Optionally, the filtering range information can be defined using text content. Specifically, the filtering range information may include starting text content and ending text content. During the filtering process, the server identifies the positions of the starting and ending text content within the image to be recognized and determines the area between their positions as the filtering range. After determining the filtering range, the server can filter out the regions to be recognized that fall within the filtering range.
[0106] Figure 6 This is a schematic diagram of the image to be identified according to an embodiment of the present invention. Figure 6 As shown, the image 611 to be recognized displays two sets of detection data. When a user wants to extract the detection time from the first detection data, the user can edit "first detection" as the starting text content and "second detection" as the ending text content into the configuration information. During the filtering process, the server can determine the positions of "first detection" and "second detection" and define the area between their positions as the filtering range, then filter out the regions 614, 612, and 615 to be recognized that are within the filtering range. In the subsequent target text extraction process, the target text will be determined in the regions 614, 612, and 615 to be recognized, thereby filtering out the region 613 to be recognized, thus preventing the device from extracting the detection time from the second detection data.
[0107] Optionally, to further improve the accuracy of the acquired target text, users can also edit multiple starting text contents and corresponding ending text contents into the configuration information. This allows the device to simultaneously filter the area to be identified within multiple filtering ranges.
[0108] It should be understood that if the configuration information does not include the filtering range information, the filtering step will be omitted. In this case, when executing step S320, the device will determine the target text region to be recognized in all regions to be recognized.
[0109] S320. Based on the text content, determine the target text region to be identified in the text region to be identified, which corresponds to each of the target basic extractors.
[0110] Specifically, after determining each of the text regions to be identified and the text content in each of the text regions to be identified, the device can determine the target text region to be identified corresponding to each of the target base extractors in the text regions to be identified based on the text content of each of the text regions to be identified.
[0111] Optionally, the server can identify the target text region corresponding to each target base extractor by determining whether the text content contains the corresponding keywords.
[0112] like Figure 5 As shown, when the target base extractor invoked by the user is the name base extractor and the gender base extractor, the device can identify the text region 53 containing the keyword "name" or "Name" and the text region 54 containing the keyword "gender" as the target text region to be identified.
[0113] Optionally, when the target text to be extracted by the target base extractor is a named entity, the server can also find the target text region to be identified corresponding to each target base extractor by determining whether the text content includes a named entity of the corresponding type.
[0114] like Figure 5 As shown, when the target base extractor called by the user is the name base extractor, the device can determine the text region 53 containing a name in the text content as the target text region to be identified.
[0115] Optionally, if there is no keyword in the target text region to be identified, the device can determine that the information of that type that the user wants to extract does not exist in the image to be identified. When providing the extraction results in the future, the device will prompt the user that the extraction of this type of information has failed.
[0116] S330. For each of the target text regions to be identified, the corresponding target basic extractor is called to extract information from the target text regions to be identified.
[0117] Specifically, after determining the target text region to be identified corresponding to each of the target basic extractors, the device can call the corresponding target basic extractor to extract information from the target text region to be identified for each target text region to be identified.
[0118] like Figure 5 As shown, for the target text region 53 to be identified, the device can call the name base extractor to extract information from the target text region 53 to determine the specific name information. For the target text region 54 to be identified, the device can call the gender base extractor to extract information from the target text region 53 to determine the specific gender information.
[0119] Optionally, each of the basic extractors can extract target text based on the text features of the target text and / or the positional relationship between the target text and the corresponding keyword. Specifically, the device can acquire the text features of the target text corresponding to the keyword and the positional relationship between the keyword and the target text, and detect the target text corresponding to the keyword in the text content of the target text region to be identified based on the text features and / or the positional relationship.
[0120] Optionally, when the target text extracted by the basic extractor is a named entity, the basic extractor can also directly extract the named entity of the corresponding type as the target text.
[0121] Optionally, during the process of determining the text region to be identified, if the keyword and the corresponding target text are not horizontally adjacent or are far apart, the keyword and the corresponding target text may be divided into two different text regions to be identified. In this case, the text region containing the keyword will not be able to extract the target text corresponding to the keyword when it is used as the target text region for information extraction. To address this, this embodiment also provides a target text extraction method, which corrects the target text region when no target text is detected, and determines the target text based on the corrected text region.
[0122] Figure 7 This is a flowchart of the target text extraction method according to an embodiment of the present invention. Figure 7 As shown, the target text extraction method may specifically include the following steps:
[0123] S400: In response to the fact that the target text is not detected in the text content of the target text region to be identified, a text region to be identified that matches the target text region to be identified is determined.
[0124] Specifically, when no target text is detected in the text content of the target text region to be identified, the device can determine a text region to be identified that matches the target text region to be identified.
[0125] Optionally, the device can determine the text region to be identified that matches the target text region to be identified based on target location information and target distance information. Here, the target location information refers to the positional information between the target text region to be identified and each of the other text regions to be identified, and the target distance information refers to the distance information between the target text region to be identified and each of the other text regions to be identified. Specifically, the device can determine the text region to be identified that is closest to the target text region to be identified in the horizontal or vertical direction as the text region to be identified that matches the target text region to be identified based on the target location information and target distance information.
[0126] Figure 8 This is a schematic diagram of the image to be identified according to an embodiment of the present invention. Figure 8As shown in the left half, the image to be recognized 81 includes a target text region 82 and a text region 83. When the device does not detect target text corresponding to the keyword "bill amount" in the target text region 82, the device can determine the text region 83, which is closest to the target text region 82 in the vertical direction, as the text region to be recognized that matches the target text region 82.
[0127] S500: Determine the corrected text region in the image to be recognized.
[0128] Specifically, after determining the text region to be recognized that matches the target text to be recognized, the device can determine a corrected text region to be recognized in the image to be recognized. The corrected text region to be recognized includes both the target text region to be recognized and the matching text region to be recognized.
[0129] like Figure 8 As shown in the right half, the device can determine a corrected text region 84 in the image to be recognized, which includes a target text region 82 and a matching text region 83.
[0130] It should be understood that, in order to facilitate the processing of the corrected text region to be recognized, the server may also include a local image preprocessor (LP_IMG) and a local text preprocessor (LP_OCR). The local image preprocessor can be used to perform image preprocessing operations on the corrected text region to be recognized, and the local text preprocessor can be used to perform corresponding text filtering operations on the corrected text region to be recognized according to user needs.
[0131] S600: Call the corresponding target base extractor to extract information from the corrected text region to be identified.
[0132] Specifically, after determining the text region to be corrected, the device can re-call the corresponding target base extractor to extract information from the corrected text region.
[0133] like Figure 8 As shown in the right half, the device can obtain the target text "xx yuan" corresponding to the keyword "bill amount" by calling the corresponding target base extractor to extract information from the corrected text region 84.
[0134] It should be understood that if the target text is still not detected, the device can return and continue to execute steps S400-S600 until the target text is determined.
[0135] Optionally, after extracting each of the target texts, the device can output each of the target texts. Specifically, when the device is a terminal, the device can display each of the target texts. When the device is a server, the device can feed back each of the target texts to the terminal, instructing the terminal to display the extraction results to the user.
[0136] Optionally, the image information extraction method in this embodiment can also be used to verify information in an image. Specifically, after obtaining the target text, the device can compare the target text with the corresponding verification text. When both meet the corresponding requirements, such as being completely identical, the device can confirm that the target text has been successfully verified. When providing the extraction result feedback, the device can provide the verification result of the target text to the user. The verification text can be specifically determined by the device in the configuration information edited by the user.
[0137] The image information extraction method of this invention, after acquiring the image to be identified and configuration information, invokes a target basic extractor to extract information from the image to be identified based on the configuration information to determine each target text in the image. The configuration information represents the invocation requirement for at least one target basic extractor, and each target basic extractor is used to extract target text of a corresponding information type. This method can reduce the cost and improve the efficiency of image information extraction.
[0138] Figure 9 This is a schematic diagram of an image information extraction device according to an embodiment of the present invention. Figure 9 As shown, the image information extraction device of this embodiment includes an image acquisition unit 91, a configuration information acquisition unit 92, and an extraction unit 93.
[0139] Specifically, the image acquisition unit 91 is used to acquire the image to be recognized;
[0140] The configuration information acquisition unit 92 is used to acquire configuration information, which is used to characterize the call requirement for at least one target basic extractor, and each target basic extractor is used to extract target text of the corresponding information type.
[0141] The extraction unit 93 is used to call each of the target base extractors to extract information from the image to be identified to determine each of the target texts in the image to be identified.
[0142] After acquiring the image to be recognized and configuration information, the image information extraction device of this embodiment of the invention calls a target basic extractor to extract information from the image to be recognized according to the configuration information to determine each target text in the image. The configuration information represents the call requirement for at least one target basic extractor, and each target basic extractor is used to extract target text of a corresponding information type. This device can reduce the cost of image information extraction and improve its efficiency.
[0143] Figure 10 This is a schematic diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a server, a terminal, etc. Figure 10 As shown, the electronic device includes at least one processor 101; a memory 102 communicatively connected to at least one processor 101; and a communication component 103 communicatively connected to a scanning device, wherein the communication component 103 receives and transmits data under the control of the processor 101; wherein the memory 102 stores instructions executable by at least one processor 101, which are executed by at least one processor 101 to implement the above-described image information extraction method.
[0144] Specifically, the electronic device includes: one or more processors 101 and a memory 102. Figure 10 Taking a processor 101 as an example, the processor 101 and the memory 102 can be connected via a bus or other means. Figure 10 Taking a bus connection as an example, memory 102, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Processor 101 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in memory 102, thereby realizing the above-mentioned image information extraction method.
[0145] Memory 102 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store an option list, etc. Furthermore, memory 102 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 102 may optionally include memory remotely located relative to processor 101, and these remote memories may be connected to external devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0146] One or more modules are stored in memory 102, and when executed by one or more processors 101, they perform the image information extraction method in any of the above method embodiments.
[0147] The above-mentioned products can perform the methods provided in the embodiments of this application, and have the corresponding functional modules and beneficial effects of performing the methods. For technical details not described in detail in this embodiment, please refer to the methods provided in the embodiments of this application.
[0148] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program for use by a computer to execute some or all of the above-described method embodiments.
[0149] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0150] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. For those skilled in the art, the present invention can be modified and varied in various ways. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for extracting image information, characterized in that, The method includes: Acquire the image to be recognized; Obtain configuration information, which is used to characterize the call requirement for at least one target basic extractor. Each target basic extractor is used to extract target text of the corresponding information type. Each target basic extractor has a unique corresponding keyword or named entity type. The configuration information includes the start text content and the end text content. Each of the aforementioned target base extractors is invoked to extract information from the image to be identified, thereby determining each of the aforementioned target texts in the image to be identified; The step of calling each of the target base extractors to extract information from the image to be identified includes: Based on the starting text content and the ending text content, locate the starting region and the ending region in the text region to be identified, and define the filtering range based on the spatial relationship between the starting region and the ending region in the image. The text regions in the image to be identified that are within the filtering range and contain the corresponding keywords or the corresponding type of named entities are determined as the target text regions to be identified. The target basic extractor is invoked to extract the target text region to be identified; In response to the fact that the target text is not detected in the text content of the target text region to be identified, a text region to be identified that matches the target text region to be identified is determined; In the image to be identified, a corrected text region to be identified is determined, the corrected text region to be identified including the target text region to be identified and the matching text region to be identified; The corresponding target base extractor is invoked to extract information from the corrected text region to be identified.
2. The method according to claim 1, characterized in that, Determining the target text region corresponding to each of the target base extractors in the text region to be identified based on the text content includes: In response to the absence of filtering range information in the configuration information, the target text region to be identified is determined in the text region to be identified based on the text content, and the target base extractor is corresponding to each of the target base extractors.
3. The method according to claim 1, characterized in that, The step of determining at least one text region to be identified in the image to be identified and the text content in each of the text regions to be identified includes: At least one text region to be identified is determined in the image to be identified; Text recognition is performed on each of the text regions to be identified to determine the text content in each of the text regions to be identified.
4. The method according to claim 3, characterized in that, Before performing text recognition on each of the text regions to be recognized to determine the text content in each of the text regions to be recognized, the method further includes: Determine the text orientation of each of the text regions to be identified; For each of the text regions to be identified, the text regions to be identified are rotated according to the text direction so that the text direction of the text regions to be identified is horizontal.
5. The method according to claim 1, characterized in that, Determining the target text region corresponding to each of the target base extractors in the text region to be identified based on the text content includes: For each of the target base extractors, the text region to be identified containing the corresponding keywords in the text content is determined as the target text region to be identified corresponding to the target base extractor.
6. The method according to claim 1, characterized in that, Determining the target text region corresponding to each of the target base extractors in the text region to be identified based on the text content includes: For each of the target base extractors, the text region to be identified that contains named entities of the corresponding type in the text content is determined as the target text region to be identified corresponding to the target base extractor.
7. The method according to claim 5, characterized in that, The step of calling the corresponding target base extractor to extract information from the target text region to be identified includes: Obtain the text features of the target text corresponding to the keywords; Obtain the positional relationship between the keywords and the target text; Based on the text features and / or the positional relationships, the target text corresponding to the keyword is detected in the text content of the target text region to be identified.
8. The method according to claim 1, characterized in that, The process of determining the text region to be identified that matches the target text region to be identified includes: The target location information and target distance information are used to determine the text region to be identified that matches the target text region to be identified; The target location information is the location information between the target text region to be identified and each of the text regions to be identified, and the target distance information is the distance information between the target text region to be identified and each of the text regions to be identified.
9. The method according to any one of claims 1-8, characterized in that, After determining each of the target texts in the image to be identified, the method further includes: Output the target text as described.
10. The method according to any one of claims 1-8, characterized in that, The configuration information is also used to characterize the invocation requirement for at least one target global image preprocessor; Before invoking each of the target base extractors to extract information from the image to be identified, the method further includes: Based on the configuration information, the target global image preprocessor is invoked to perform image preprocessing on the image to be identified.
11. The method according to any one of claims 1-8, characterized in that, The method further includes: In response to receiving an information extraction request, the image to be identified and the configuration information are obtained according to the information extraction request.
12. An image information extraction device, characterized in that, The device includes: An image acquisition unit is used to acquire the image to be recognized. A configuration information acquisition unit is used to acquire configuration information, which is used to characterize the call requirement for at least one target basic extractor. Each target basic extractor is used to extract target text of the corresponding information type. Each target basic extractor has a unique corresponding keyword or named entity type. The configuration information includes filtering range information, which includes starting text content and ending text content. An extraction unit is used to call each of the target basic extractors to extract information from the image to be identified to determine each of the target texts in the image to be identified. The extraction unit is also used for: Based on the starting text content and the ending text content, locate the starting region and the ending region in the text region to be identified, and define the filtering range based on the spatial relationship between the starting region and the ending region in the image. The text regions in the image to be identified that are within the filtering range and contain the corresponding keywords or the corresponding type of named entities are determined as the target text regions to be identified. The target basic extractor is invoked to extract the target text region to be identified; In response to the fact that the target text is not detected in the text content of the target text region to be identified, a text region to be identified that matches the target text region to be identified is determined; In the image to be identified, a corrected text region to be identified is determined, the corrected text region to be identified including the target text region to be identified and the matching text region to be identified; The corresponding target base extractor is invoked to extract information from the corrected text region to be identified.
13. A computer-readable storage medium storing computer program instructions thereon, characterized in that, The computer program instructions, when executed by a processor, implement the method as described in any one of claims 1-11.
14. An electronic device, characterized in that, The device includes: Memory is used to store one or more computer program instructions; A processor, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of claims 1-11.
Citation Information
Patent Citations
Data element profiles and overrides for dynamic optical character recognition based data extraction
US10740638B1
Information representation structure analysis device, and information representation structure analysis method
WO2022215433A1