Domain-based Text Extraction Method and System

The method addresses the challenge of extracting information from complex documents by generating search patterns from user input and associating them with text entity patterns, resulting in effective extraction and representation of relevant data.

JP7696893B2Active Publication Date: 2025-06-23L&T TECH SERVICES LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022525481
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-30
Filing Date
2020-12-30
Publication Date
2025-06-23
Estimated Expiration
2040-12-30

AI Technical Summary

Technical Problem

Existing techniques struggle to extract relevant information from complex documents without tags or identifiable patterns, making it difficult to generate accurate representations of the extracted data.

Method used

The method involves identifying text data from an input file, receiving user input to identify related text entities, automatically generating a search pattern, associating it with patterns from the text entities, and extracting matching patterns to generate a representation of the extracted information.

Benefits of technology

This approach enables effective extraction of relevant information even from complex documents without tags or patterns, allowing for accurate generation of representations based on the extracted data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696893000001
    Figure 0007696893000001
  • Figure 0007696893000002
    Figure 0007696893000002
  • Figure 0007696893000003
    Figure 0007696893000003
Patent Text Reader

Abstract

The present disclosure relates to a method and system for extracting information from the contents of an input file. The method may include identifying text data from the input file, receiving text input from a user to identify related text entities from a plurality of text entities, and automatically generating a search pattern corresponding to the text input. The method may further include determining a pattern associated with each of the plurality of text entities and matching the search pattern corresponding to the text input with the patterns associated with the plurality of text entities. The method may further include identifying one or more matching patterns from the patterns associated with the plurality of text entities based on the matching, and extracting related text entities from the plurality of text entities that correspond to the one or more matching patterns.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to data extraction, and more particularly, to methods and systems for extracting text information from the content of an input file using one or more data extraction approaches.

Background Art

[0002] Recently, text extraction techniques have become important. For example, extraction techniques such as Optical Character Recognition (OCR) can enable a user to extract text data from files such as images or Portable Document Format (PDF) files. Further, it may be desirable to extract relevant information and generate a representation using the extracted relevant information.

Summary of the Invention

Problems to be Solved by the Invention

[0003] According to some available techniques, it may be possible to extract text information from an input file and generate a representation using the extracted information if the input file contains tags or if patterns can be identified in the identified text. However, in complex documents where tags may be absent or patterns cannot be identified, it is difficult to extract relevant information or generate a representation.

Means for Solving the Problems

[0004] The method and system according to the present invention extract information from the content of an input file, and include identifying text data from the input file, receiving text input from a user to identify related text entities from a plurality of text entities, automatically generating a search pattern corresponding to the text input, determining a pattern associated with each of the plurality of text entities, associating the search pattern corresponding to the text input with the pattern associated with the plurality of text entities, identifying one or more matching patterns from the patterns associated with the plurality of text entities based on the association, and extracting related text entities corresponding to the one or more matching patterns from the plurality of text entities. The identified text data includes a plurality of text entities.

[0005] In addition, the problems disclosed in the present application and the solutions thereto are clarified by the description in the section of the mode for carrying out the invention and the drawings.

[0006] The accompanying drawings incorporated in and constituting a part of this disclosure illustrate exemplary embodiments and serve to explain the principles of the disclosure together with the description.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0008] Exemplary embodiments will be described with reference to the accompanying drawings. Where convenient, the same reference numbers are used throughout the drawings to indicate the same or similar parts. Examples and features of the principles of the disclosure are described herein, but modifications, alterations, and other implementations are possible without departing from the spirit and scope of the embodiments of the disclosure. The following detailed description is intended to be regarded as merely exemplary, and the true scope and spirit are indicated by the following claims.

[0009] Referring now to FIG. 1, there is shown a functional block diagram of an exemplary system 100 for extracting information from the content of an input file according to some embodiments of the present disclosure. System 100 may include a text identification device 102. In some embodiments, the text identification device 102 may identify text data from the content of the input file. For example, the input file may be an image or a Portable Document Format (PDF) file. The identified text data may include a plurality of text entities. Thus, in some embodiments, the text identification device 102 may use Optical Character Recognition (OCR) techniques to identify text data from the input file. In alternative embodiments, the text identification device 102 may use any other techniques known in the art to identify text data from the input file.

[0010] System 100 may further include an information extraction device 104 communicable with the text identification device 102. For example, the information extraction device 104 may be configured to extract relevant information from a plurality of text entities identified from the input file by the text identification device 102. In some embodiments, to extract information, the information extraction device 104 may receive text input from a user to identify relevant text entities from the plurality of text entities, and automatically generate a search pattern corresponding to the text input. The information extraction device 104 may further determine a pattern associated with each of the plurality of text entities, and associate the search pattern corresponding to the text input with the pattern associated with the plurality of text entities. The information extraction device 104 may further identify one or more matching patterns from the patterns associated with the plurality of text entities based on the association, and extract relevant text entities corresponding to the one or more matching patterns from the plurality of text entities.

[0011] System 100 may further include a display 114. In some embodiments, the information extraction device 104 may include one or more processors 110 and a computer-readable medium (e.g., memory) 112. The computer-readable storage medium 112 may store instructions that, when executed by one or more processors 110, cause the one or more processors 110 to generate a representation based on the content of the input file, according to aspects of the present disclosure. The computer-readable storage medium 112 may also store various data (e.g., identified text data, value entities, data, pattern entity data, etc.) that may be captured, processed, and / or required by the system 100. The system 100 may interact with a user via a user interface 116 accessible via the display 114. The system 100 may also interact with one or more external devices 106 via a communication network 108 to send or receive various data. The external device 106 may include, but is not limited to, a remote server, a digital device, or another computing system.

[0012] Referring now to FIG. 2, a functional block diagram of a natural language processing (NLP) metadata extraction framework 200 according to some embodiments of the present disclosure is shown. The NLP metadata extraction framework 200 may be implemented in the system 100, particularly in the information extraction device 104. The NLP metadata extraction framework 200 may include various modules that perform various functions for extracting information from the content of an input file. In some embodiments, an extraction engine 202 may extract / identify text data from the content of the input file 204. The input file 204 may be a flat file 206, or a PDF file 208, or a database file 210, or an OCR output 212.

[0013] The extraction engine 202 may further receive one or more domain rules from the domain rule database 214. Further, in some embodiments, the domain rule database 214 may be based on "if-then" rules (rather than conventional procedural code). Note that the domain rules may be used to determine domain-based names associated with each of a plurality of text entities. The domain-based names may be determined based on associating each of the plurality of text entities with a list of domain-based names stored in a data depository (not shown in FIG. 1). For example, if the identified data includes a text entity belonging to a PIN code of a city in a certain country, the domain-based database may determine the domain name, i.e., the name of the city having that PIN code, by associating the PIN code with a list of domain-based names stored in the data repository.

[0014] The extraction engine 202 may further receive one or more location rules from the location rule database 216. Note that in some embodiments, the location rules may be used to determine the locations associated with entities from text data in an input file based on the OCR output. The location may be determined when the type associated with the input file is first identified and the locations of one or more related text entities in the input file are determined based on the type associated with the input file. Once the locations of the related text entities are determined, the related text entities may be extracted from the input file based on the locations. One or more location rules may be more relevant in the case of input data including a table, i.e., input data arranged in tabular form, but note that one or more location rules may be equally applicable to any other type of input data, such as a document having free text.

[0015] As an example, the type of the input file may be that of an "Aadhar" card. It should be further understood that in a document such as an "Aadhar" card, the data may exist in a specific format. Thus, the positions of the relevant text entities may be predetermined based on a template corresponding to the type associated with the input file stored in the data repository. For example, the positions of the text entities "name" and "address" (of the person in question) may be in the middle of the document, and this information may be retrievable and stored in the database of system 100. In particular, the positions may be based on (x, y) coordinates. In the above example, the positions of the text entities "name" and "address" may be within the region defined by the (x, y) coordinates (5, 10) to (15, 15). Thus, when the type of the document is identified as an "Aadhar" card, system 100 may access the above positions and extract the relevant text entities (i.e., "name" and "address") from those positions.

[0016] The extraction engine 202 may further receive one or more first NLP rules from the first NLP rule database 218. It should be noted that the first NLP rule database 218 may create a regular expression (Regex) based on the NLP rules provided for the correct values / attributes. As will be understood by those skilled in the art, a Regex may be a special string for describing a search pattern.

[0017] The extraction engine 202 may further receive one or more second NLP rules from the second NLP rule database 220. Note that in some embodiments, the second NLP rule database 220 may define a part of speech (POS) for the correct value / attribute. For example, the POS associated with a plurality of text entities may be identified. As will be understood by those skilled in the art, the POS may be one of a noun, pronoun, verb, adverb, adjective, conjunction, preposition, and interjection. When the POS is identified, one or more unique text entities may be selected from the plurality of text entities for which the POS is a noun. One or more (selected) unique text entities identified as nouns may be extracted as relevant data.

[0018] For example, a salary statement (input data) may include a plurality of text entities including the name of the person of interest. Extracting the name may be the purpose of the extraction engine 202. Further, it will be understood that the name belongs to the POS category of nouns. Thus, the text entity for which the POS is a noun will be selected. Thus, the required name will be extracted from the selected entity.

[0019] In some embodiments, when an attempt is made to identify the POS associated with a plurality of text entities, one or more unique text entities for which the POS is not identified may be identified. An entity type associated with each of the one or more unique text entities may be determined. The entity type may be a value entity or a pattern entity. Each value entity may have an associated value, and each pattern entity may have an associated pattern.

[0020] For example, an identified text entity may be composed of a combination of text characters that indicate a codename such as an email ID or an oilfield name (e.g., "john@gmail.com" or "AB-123" respectively). Thus, for such a text entity, a POS may not be identified. Thus, such a text entity may be identified as either a value entity or a pattern entity. Thus, an email ID ("john@gmail.com") may be determined as a value entity, and the value is the email ID. This may be because the pattern of an email ID is usually a typical and known pattern, and all email IDs may share the same or a similar pattern. On the other hand, a text entity such as "AB-123" corresponding to an oilfield name may be determined as a pattern entity. This is because oilfield names may share the same pattern, but such a pattern may not be generally known, i.e., the NLP metadata extraction framework 200 may not have pre-stored such a pattern. When a value entity is identified, text entities in the input file having the same associated value as the associated value of the value entity may be automatically identified. In other words, all email ID text entities may be identified and extracted.

[0021] When a pattern entity is identified, a search pattern corresponding to the pattern (Regex) associated with the pattern entity may be automatically generated. Further, the pattern associated with each of a plurality of text entities in the input file may be determined. The search pattern (Regex) corresponding to the pattern entity may be associated with the patterns associated with the plurality of text entities, and based on the association, one or more matching patterns may be determined from the patterns associated with the plurality of text entities. Matching text entities corresponding to the one or more matching patterns may be extracted from the plurality of text entities.

[0022] For example, for the text entity "AB-123" corresponding to an oil field, a search pattern (Regex) "apd" (where "a" means one or more alphabets, "p" means one or more punctuation marks, and "d" means one or more digits) may be automatically generated. Further, patterns associated with other text entities in the input file are determined, and the search pattern "apd" may be associated with the patterns associated with other text entities, and matching patterns may be identified. Thereafter, all text entities corresponding to the matching patterns are extracted. Thereby, all oil field names in the input file may be extracted.

[0023] In some embodiments, when text data from an input file is identified, the first NLP rule may be caused to receive text input from a user to identify related text entities from a plurality of text entities and automatically generate a search pattern corresponding to the text input. The first NLP rule may further be caused to determine patterns associated with each of the plurality of text entities and associate the search pattern corresponding to the text input with the patterns associated with the plurality of text entities. The first NLP rule may further be caused to identify one or more matching patterns from the patterns associated with the plurality of text entities based on the association and extract related text entities corresponding to the one or more matching patterns from the plurality of text entities.

[0024] For example, a user may provide a text input "AB-123" to identify related text entities from a plurality of text entities extracted from an input file. A search pattern (Regex) "apd" may be identified for the input text entity "AB-123". Further, patterns associated with other text entities in the input file may be determined, and the search pattern (Regex) "apd" may be associated with the patterns associated with other text entities, and matching patterns may be identified. Thereafter, all text entities corresponding to the matching patterns may be extracted. Thereby, all oil field names in the input file may be extracted.

[0025] Note that in some embodiments, the text input from the user may be received in the form of an entry in a "Microsoft Excel (registered trademark)" sheet. One or more rules may be defined that can automatically generate a search pattern (Regex) corresponding to the text input. Thereafter, the one or more rules may determine the patterns associated with each of the plurality of text entities of the input file, associate the search pattern corresponding to the text input with the patterns associated with the plurality of text entities, identify one or more matching patterns from the patterns associated with the plurality of text entities based on the association, and extract related text entities corresponding to the one or more matching patterns from the plurality of text entities.

[0026] Thus, in some embodiments, the "Microsoft Excel (registered trademark)" data may be received from one or more "Microsoft Excel (registered trademark)" sheets 222. By applying various rules (domain rules, NLP rules, location rules, etc.), the extracted data 226 may be obtained. In some embodiments, a "Microsoft Excel (registered trademark)" macro may be used for extraction. Further, in some embodiments, the macro may be embedded in Python (language; registered trademark). Here, the extracted data 226 may refer to related text entities extracted from a plurality of text entities, or expressions generated based on related text entities extracted using various rules. In some additional embodiments, if the above approaches, namely the domain-based, location-based, POS-based, and Regex-based approaches, fail to provide text extraction, a machine learning (ML)-based approach may be used. An AI model (i.e., an ML model) 224 may be used for the ML-based approach.

[0027] Referring now to FIG. 3, a block diagram of a spell-checking system 300 according to some embodiments of the present disclosure is shown. In some embodiments, text entities extracted by an extraction module 302 (corresponding to the text identification device 102) may be received by the spell-checking system 300. These expressions generated by the extraction module 302 may serve as input text 304 to the spell-checking system 300. The input text 304 may be preprocessed by a preprocessing module 306.

[0028] When preprocessing is performed, cosine similarity analysis may be performed on the preprocessed data by the cosine similarity module 308. As an example, the cosine similarity module 308 may perform cosine similarity analysis using a threshold. The minimum edit distance module 310 may perform minimum edit distance analysis on the data received from the cosine similarity module 308. In some embodiments, the minimum edit distance analysis may be performed on the filtered words.

[0029] The maximum cosine similarity module 312 may perform maximum cosine similarity analysis on the data received from the minimum edit distance module 310. It should be noted that the maximum cosine similarity analysis may be based on the words within the minimum edit distance word. The replacement module 314 may replace the inaccurate words with the corrected words based on the analysis performed by the maximum cosine similarity module 312.

[0030] Referring now to FIG. 4, a block diagram of a recommendation system 400 according to some embodiments of the present disclosure is shown. For example, the recommendation system 400 may be used to identify an entity type, i.e., a value entity or a pattern entity. In some embodiments, the recommendation system 400 may use a classification-based machine learning (ML) model 402. The recommendation system 400 may receive input data (NLP extraction / spell check data) 406. It should be noted that the input data may be data on which NLP and spell check have been performed. The recommendation system 400 may include a config file 408. The config file 408 may trigger a determination as to whether a value can be recommended for the input data or whether a pattern can be recommended.

[0031] Recommendation system 400 may supply input data to ML model 402. ML model 402 may use historical data from archive database 404. Based on the historical data, ML model 402 may provide prediction data 410 or recommendation data 412. Accordingly, a recommended value 414 may be generated based on prediction data 410, or a recommendation pattern 416 may be generated based on recommendation data 412.

[0032] Referring now to FIG. 5, a block diagram of a metadata update system 500 according to some embodiments of the present disclosure is shown. In some embodiments, metadata update system 500 may receive NLP-extracted data 502. Metadata update system 500 may include a system date / time module 504. Metadata update system 500 may further include a confidence module 506 and a probability module 508. Confidence module 506 may determine a confidence score associated with the accuracy of the extracted data and the expressions generated by the NLP rules. Probability module 508 may determine an accuracy score associated with the accuracy of the extracted data and the expressions generated by the NLP rules based on the confidence score. In other words, probability module 508 may determine the probability of how accurate / useful the expressions generated using the extracted data or NLP rules are.

[0033] If the probability determined by probability module 508 is high enough (i.e., greater than a threshold), an expression generated using the extracted data and NLP rules, i.e., a rule-based value 514, may be provided. On the other hand, if the probability determined by probability module 508 is not high enough (i.e., less than a threshold), an expression generated using the extracted data and an ML model, i.e., a machine learning (ML) value 516, may be provided. As described above, the ML value may be generated by an ML model using historical data that may be stored in archive database 510. The extracted data and the generated expressions may be stored in final database 512.

[0034] Referring now to FIG. 6, a flowchart of a method 600 for extracting information from the content of an input file according to an embodiment of the present disclosure is shown. As will be appreciated, method 600 may provide a Regex-based method of text extraction. At step 602, text data may be identified from the input file. The identified text data may include a plurality of text entities. Note that in some examples, the input file may include at least one of an image file and a Portable Document Format (PDF) file.

[0035] At step 604, text input may be received from a user to identify related text entities from the plurality of text entities. For example, the user may provide the text as an entry in a "Microsoft Excel (registered trademark)" sheet. At step 606, a search pattern (Regex) corresponding to the text input may be automatically generated. At step 608, a pattern associated with each of the plurality of text entities may be determined. At step 610, the search pattern corresponding to the text input may be associated with the patterns associated with the plurality of text entities. At step 612, one or more matching patterns from the patterns associated with the plurality of text entities may be identified based on the association. At step 614, related text entities corresponding to the one or more matching patterns may be extracted from the plurality of text entities.

[0036] In addition, in some embodiments, when extracting related text entities corresponding to one or more matching patterns from the plurality of text entities, an expression may be generated using the extracted related text entities.

[0037] Furthermore, for example, in method 600, a machine learning (ML) model may be used to extract related entities if they cannot be extracted. In other words, if data (related text entities) cannot be extracted, an ML-based approach may be initiated.

[0038] Referring now to FIG. 7, there is shown a flowchart of a method 700 for extracting information from the content of an input file, according to another embodiment of the present disclosure. At step 702, text data may be identified from the input file. The identified text data may include a plurality of text entities. Note that the input file may include at least one of an image file and a Portable Document Format (PDF) file.

[0039] At step 704, a domain-based name associated with each of the plurality of text entities may be determined. The domain-based name may be determined based on associating each of the plurality of text entities with a list of domain-based names stored in a data repository. At step 706, a type associated with the input file may be identified. At step 708, based on the type associated with the input file, the location of one or more related text entities in the input file may be determined. At step 710, based on the location, one or more related text entities may be extracted from the input file.

[0040] At step 712, a part of speech (POS) associated with the plurality of text entities may be identified. Understand that the POS may be one of a noun, pronoun, verb, adverb, adjective, conjunction, preposition, and interjection. At step 714, one or more text entities identified as nouns may be determined.

[0041] In step 716, one or more unique text entities for which the POS is not identified may be selected. In step 718, the entity type associated with each of the one or more unique text entities may be determined. The entity type may be one of a value entity and a pattern entity. Further, each value entity may have an associated value, and each pattern entity may have an associated pattern. As described in connection with FIG. 4, a (ML-based) recommendation system 400 may be used to determine the entity type associated with each of the one or more unique text entities.

[0042] In step 720, for each value entity, one or more text entities having the same associated value as the associated value of the value entity may be automatically identified. For each pattern entity, steps 722-728 may be performed. Thus, in step 722, a search pattern (Regex) corresponding to the pattern associated with the pattern entity may be automatically generated. In step 724, the search pattern corresponding to the pattern entity may be associated with the patterns associated with the plurality of text entities. In step 726, based on the association, one or more matching patterns may be determined from the patterns associated with the plurality of text entities. In step 728, the matching text entities corresponding to the one or more matching patterns may be extracted from the plurality of text entities.

[0043] Note that method 700 may be performed in combination with method 600. In an exemplary scenario, method 600 may start after step 710 of method 700 is completed and before step 712 of method 700 begins.

[0044] Referring now to FIG. 8, there is shown a flowchart of a method 800 for extracting information from the content of an input file, according to another embodiment of the present disclosure. In step 802, text data may be extracted from the content of the input file. The extracted text data may include a plurality of text entities.

[0045] In step 804, a domain-based name associated with each of the plurality of text entities may be determined. The domain-based name may be determined based on associating each of the plurality of text entities with a list of domain-based names stored in a data repository. For example, if a PIN code is obtainable, the city may be determined from the PIN code using a domain-based approach by, for example, associating the PIN code with a list (of PIN codes and associated city names) stored in a knowledge base. Similarly, if a city name (e.g., Delhi) is identified, the country name (e.g., India) may be determined. In some embodiments, accordingly, a tag may be assigned to each of a second plurality of entities.

[0046] In step 806, a check may be made to determine whether the determination of the domain-based name was successful. If the determination of the domain-based name was successful, the method may proceed to step 822 ("Yes" path), where the identified text entities may be extracted. Further, a representation may be generated using the extracted text entities. In step 806, if the domain-based name is not determined, the method may proceed to step 808 ("No" path). In other words, if the domain-based name cannot be identified, the method may proceed to the execution of the next alternative approach.

[0047] In step 808, based on the positions of each of the plurality of text entities, the value of a predetermined field may be determined. The positions may be in the form of (x and y) coordinates. As an example, the positions may be predetermined based on a template stored in a data repository. For example, in some scenarios, the value of a particular name may be described next to that name or below a particular name header, which may be identifiable based on (x and y) coordinates. For example, in an "Aadhar card", or "PAN card", or driving license, an entity, such as the name of a user, may be obtainable at a particular position from which the name can be identified. The technique of positions may be particularly useful in the case of a table. This is because the data in the table is available in a structured form and the positions of the entities can be easily and accurately identified. Thus, the type associated with the input file, i.e., whether the input file is an "Aadhar card", or a "PAN card", or a driving license, may first be identified. Further, based on the type associated with the input file, the positions of one or more relevant text entities in the input file may be determined.

[0048] In step 810, a check may be made to confirm whether or not the determination of the positions of one or more relevant text entities in the input file has been successful. If the determination of the positions is successful, method 800 may proceed again to step 822 ("Yes" path), at which time one or more relevant text entities from the input file may be extracted based on the positions and a representation may be generated using the extracted text entities. If it is determined that the positions are not determined, in step 810, method 800 may proceed to step 812.

[0049] In step 812, to extract related text entities, a Regex-based approach may be tried on the input file. For this purpose, text input may be received from the user to identify related text entities from a plurality of text entities, and a search pattern corresponding to the text input may be automatically generated. In other words, a Regex may be generated. Further, patterns associated with each of the plurality of text entities may be generated, and the search pattern (Regex) corresponding to the text input may be associated with the patterns associated with the plurality of text entities, and based on the association, one or more matching patterns may be identified from the patterns associated with the plurality of text entities. In some embodiments, method 800 may include calculating a probability of determining related entities from the extracted text data corresponding to the Regex.

[0050] In step 814, a check may be made to determine whether the Regex-based approach has been successfully applied. If the Regex-based approach has been successfully applied, method 800 may proceed to step 822 (the "Yes" path), at which time related text entities corresponding to one or more matching patterns may be extracted from the plurality of text entities, and an expression may be generated using the extracted text entities. If the Regex-based approach is not successful, method 800 may proceed to step 816 (the "No" path).

[0051] In step 816, a part-of-speech (POS)-based approach may be tried. For example, the POS associated with each of a plurality of text entities may be determined. In other words, in step 816, a trial may be conducted to determine with which POS an entity can be associated, for example, whether the POS can be a noun, an adverb, an adjective, etc. Further, in step 816, a trial may be conducted to determine the entity type associated with each of one or more unique text entities. For this purpose, one or more unique text entities for which the POS is not identified may be selected. For example, for a unique entity having a combination of alphabets, numbers, punctuation marks, etc. (such as the email ID or oil field name described above), the POS may not be identified. For such a text entity, the entity type associated with each of one or more unique text entities may be determined. The entity type may be one of a value entity and a pattern entity. Each value entity may have an associated value, and each pattern entity may have an associated pattern. Further, for each value entity, a trial may be conducted to automatically identify one or more text entities having the same associated value as the associated value of the value entity (for example, the email ID).

[0052] In step 818, a check may be made to determine whether the trial was successful (i.e., whether the POS-based approach was successful). If the trial is determined to be successful, method 800 may proceed to step 822 (the "Yes" path), at which time one or more text entities (associated text entities) having the same associated value as the associated value of the value entity may be extracted, and an expression may be generated using the extracted text entities.

[0053] Furthermore, for each pattern entity, a search pattern corresponding to the pattern associated with the pattern entity may be automatically generated, the search pattern corresponding to the pattern entity may be associated with the patterns associated with a plurality of text entities, and a trial may be performed to determine one or more matching patterns from the patterns associated with the plurality of text entities based on the association. In step 818, a check may be performed to determine whether the trial was successful (i.e., whether the POS-based approach was successful). If the trial is determined to be successful, method 800 may proceed to step 822 ("Yes" path), at which time one or more matching text entities corresponding to the one or more matching patterns may be extracted from the plurality of text entities, and an expression may be generated using the extracted text entities. In step 818, if the POS-based approach is determined not to be successful, method 800 may proceed to step 820 ("No" path).

[0054] In step 820, a trial may be performed to extract relevant text entities (information) from a plurality of text entities using a machine learning (ML) model. It will be appreciated that the ML model may first be trained to extract the type of relevant text entity. Thus, based on the training of the ML model, an ML-based classification may be applied to the input file to extract relevant text entities. Thus, in step 822, relevant text entities may be extracted based on the ML-based classification, and an expression may be generated using the extracted text entities.

[0055] Note that in some embodiments, the various steps (i.e., steps 804, 808, 812, 816, and 820) may be performed in the order described above or in any other order.

[0056] Furthermore, note that an NLP-based approach may be used in the following scenario. (i) When a search key (i.e., text input from the user) is available and text extraction (i.e., related entity) associated therewith is expected in the input file. (ii) When a search key (i.e., text input from the user) is not available but text extraction (i.e., related entity) associated therewith is expected in the input file. (iii) When a search key (i.e., text input from the user) is available but text extraction (i.e., related entity) associated therewith is not expected in the input file.

[0057] It should be noted that a confidence score (probability score) may be calculated in each approach. For example, for the "domain-based" approach, a maximum probability score (e.g., 1) may be obtained when the search key (i.e., text input from the user) exactly matches the remaining text entities in the input file. Further, in the "location-based" approach, the probability score (p) may be based on the probability (w) of the search key match and the identification of the location (l), i.e., p = w * l. For the Regex-based approach, the probability for location identification may be either 0 or 1, and the probability score (e) for extraction may be 1 or 0. Further, the values to be extracted may be multiple (n), and one of those values may be selected. Therefore, the probability of the selected value may be "1 / n", and the confidence score may be e * (1 / n). When both the "location-based" and "Regex-based" approaches are used, a weight of 50% may be given to each approach. Therefore, the confidence score may be (w * l) / 2 + (e * (1 / n)) / 2. For the ML-based approach, the probability score may be derived based on the confusion matrix (accuracy rate).

[0058] Furthermore, when the following approaches, namely a pure word matching approach, a rule-based approach (i.e., one or more of a domain-based, position-based, POS-based, and Regex-based approach), and an ML-based approach are used, a composite probability score (P) may be calculated. Therefore, Composite probability score (P) = (P1 * 0.3) * (P3 * 0.7), or Composite probability score (P) = (P1 * 0.3) * (P2 * 0.7) wherein, P1 = probability score of the pure word matching approach, P2 = probability score of the rule-based approach, and P3 = probability score of the ML-based approach.

[0059] <Case scenario 1> Referring now to FIG. 9, a snapshot of an exemplary input file 900 is shown where a noisy entity 902 exists in the input file 900. For example, the first part of the entity may be illegible due to smudging or fading, etc., and the second part of the entity may contain the word "United". Thus, the actual entity may be "United Airlines (registered trademark)" or "United States", etc. To determine this entity, one or more approaches (i.e., domain-based, position-based, POS-based, Regex-based, or ML-based) may be used to determine the actual entity. For example, if the domain-based, position-based, POS-based, and Regex-based approaches fail to provide an extraction, the ML-based approach may be initiated and text extraction may be provided based on historical data using a classification-based machine learning (ML) model. The ML model may be a deterministic model and / or a probabilistic model.

[0060] <Case scenario 2> Referring now to FIG. 10, a snapshot of the input file "Local Order" 1000 is shown, from which various text entities may need to be extracted. As shown in FIG. 10, the input file "Local Order" 1000 has a tabular structure.

[0061] For example, in order to extract the value for the attribute "RO Number", the "RO Number" may be extracted from the domain dictionary database (lookup table) 1002 by associating the RO date attribute.

[0062] To extract the attribute Part Number, the key "Part Number" existing in the input file 1000 may be identified or received as user input. Thereafter, using the template information stored in the database by a position-based approach, the possible position of the Part Number may be determined. For example, by a position-based approach, it may be suggested that the Part Number may exist on the "right" side of the input file 1000. Therefore, the required Part Number "32145643" may be extracted from the cell coordinates on the right of the tabular structure of the input file "Local Order" 1000.

[0063] To extract the attributes of the email IDs present in the input file 1000, a Regex-based approach may be used. Thus, the keys (Regex) associated with the attributes of the email may be identified or received from a database. Thus, the Regex may be "apapapa" (where "a" means one or more alphabets, "p" means one or more punctuation marks, and "d" means one or more digits) corresponding to the email ID "ajay.thakur@gmail.com". Thus, the text entity having a matching pattern (Regex), i.e., the text entity of the email ID, may be identified and the value "ajay.thakur@gmail.com" may be extracted from the corresponding cell of the tabular structure of the input file 1000.

[0064] To extract the attribute Name from the input file 1000, a POS-based approach may be used. For example, first, the key present for the attribute (Name) may be identified in the input file 1000. Thereafter, for example, using a Named Entity Recognizer (NER) tagger, the POS of the identified text entity in the input file 1000 may be determined. The NER tagger specified in the database for the attribute Name may be searched. Thus, the text entity corresponding to the attribute Name ("Ajay Thakur") having the POS of a noun may be extracted. Thus, an expression using the extracted text entity "Ajay Thakur" may be generated.

[0065] One or more text extraction techniques for extracting information from the content of an input file are disclosed above. The above techniques provide one or more advantages over conventional techniques. For example, these techniques provide various approaches for text extraction, namely domain-based, location-based, POS-based, Regex-based approaches, and ML-based approaches. Thus, text extraction can be performed using a single approach, or multiple approaches applied in any order. Further, these techniques enable the generation of representations using the extracted information. Additionally, these techniques enable the extraction of information and the generation of representations using the extracted information even in complex documents where there are no tags or patterns cannot be identified in the extracted information.

[0066] In implementations of embodiments consistent with the present disclosure, one or more computer-readable storage media may be utilized. A computer-readable storage medium refers to any type of physical memory where information or data readable by a processor can be stored. Thus, a computer-readable storage medium may store instructions for one or more processors to execute steps or stages consistent with the embodiments described herein. The term "computer-readable medium" is to be understood to include tangible articles and to exclude carrier waves and transient signals, i.e., it is non-transitory. Examples include random access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, CD-ROMs, DVDs, flash drives, disks, and any other known physical storage media.

[0067] The disclosures and examples are intended to be regarded as merely illustrative, and the true scope and spirit of the embodiments of the disclosure are indicated by the following claims.

Claims

1. A method for extracting information from the content of an input file, comprising: a computer identifying text data from the input file, the identified text data including a plurality of text entities; receiving text input from a user to identify related text entities from the plurality of text entities; automatically generating a search pattern corresponding to the text input; determining a pattern associated with each of the plurality of text entities; associating the search pattern corresponding to the text input with the patterns associated with the plurality of text entities; identifying one or more matching patterns from the patterns associated with the plurality of text entities based on the association; extracting related text entities corresponding to the one or more matching patterns from the plurality of text entities and a method including the above.

2. The method according to claim 1, further comprising generating a representation using the extracted related text entities.

3. The method according to claim 1, wherein the input file includes at least one of an image file and a Portable Document Format (PDF) file.

4. The method according to claim 1, further comprising determining a domain-based name associated with each of the plurality of text entities, wherein the domain-based name is determined based on associating each of the plurality of text entities with a list of domain-based names stored in a data repository. The method according to claim 1.

5. Identifying the type associated with the input file; Determining the positions of one or more related text entities in the input file based on the type associated with the input file; Extracting the one or more related text entities from the input file based on the positions; The method according to claim 1, further comprising.

6. The method according to claim 5, wherein the position is based on (x, y) coordinates and is predetermined based on a template corresponding to the type associated with the input file stored in the data repository.

7. Identifying the part-of-speech (POS) associated with the plurality of text entities; Extracting one or more text entities for which the POS is identified as a noun; Further comprising, wherein the POS is one of a noun, pronoun, verb, adverb, adjective, conjunction, preposition, and interjection. The method according to claim 1.

8. Identifying one or more unique text entities for which the POS is not identified; Determining the entity type associated with each of the one or more unique text entities; Further comprising, wherein the entity type is one of a value entity and a pattern entity, each value entity has an associated value, and each pattern entity has an associated pattern. The method according to claim 7.

9. For each value entity, Automatically identifying one or more text entities having the same associated value as the associated value of the value entity. The method according to claim 8, further comprising

10. For each pattern entity, automatically generating a search pattern corresponding to the pattern associated with the pattern entity; associating the search pattern corresponding to the pattern entity with a pattern associated with the plurality of text entities; determining one or more matching patterns from the patterns associated with the plurality of text entities based on the association; extracting matching text entities corresponding to the one or more matching patterns from the plurality of text entities, and The method according to claim 8, further comprising

11. The method according to claim 1, comprising generating a recommendation based on historical data using a classification-based machine learning (ML) model.

12. A system for extracting information from the content of an input file, comprising a processor; a memory communicatively coupled to the processor, and wherein the memory stores processor-executable instructions that, when executed by the processor, cause the processor to identify text data from the input file, the identified text data including a plurality of text entities; receive text input from a user to identify related text entities from the plurality of text entities; automatically generate a search pattern corresponding to the text input; determine a pattern associated with each of the plurality of text entities; determine a pattern associated with each of the plurality of text entities; Associating the search pattern corresponding to the text input with the patterns associated with the plurality of text entities. Identifying one or more matching patterns from the patterns associated with the plurality of text entities based on the association, and Causing the related text entities corresponding to the one or more matching patterns to be extracted from the plurality of text entities. System.

13. The system according to claim 12, wherein the input file includes at least one of an image file and a Portable Document Format (PDF) file.

14. The processor-executable instructions, when executed by the processor, cause the processor to Perform domain-based extraction, where performing the domain-based extraction Determine a domain-based name associated with each of the plurality of text entities, where the domain-based name is determined based on associating each of the plurality of text entities with a list of domain-based names stored in a data repository. Including domain-based extraction. Perform location-based extraction, where performing the location-based extraction Identify the type associated with the input file, Determine the location of one or more related text entities in the input file based on the type associated with the input file, and Extract the one or more related text entities from the input file based on the location, where the location is predetermined based on a template corresponding to the type associated with the input file stored in a data repository. Including location-based extraction. Part-of-speech (POS)-based extraction, wherein performing the POS-based extraction comprises identifying POS associated with the plurality of text entities, wherein the POS is one of a noun, pronoun, verb, adverb, adjective, conjunction, preposition, and interjection, and extracting one or more text entities for which the POS is identified as a noun, and identifying one or more unique text entities for which the POS is not identified, and determining an entity type associated with each of the one or more unique text entities, wherein the entity type is one of a value entity and a pattern entity, each value entity has an associated value, and each pattern entity has an associated pattern, and for each value entity automatically identifying one or more text entities having the same associated value as the associated value of the value entity, and for each pattern entity automatically generating a search pattern corresponding to the pattern associated with the pattern entity, and associating the search pattern corresponding to the pattern entity with patterns associated with the plurality of text entities, and determining one or more matching patterns from the patterns associated with the plurality of text entities based on the association, and extracting matching text entities corresponding to the one or more matching patterns from the plurality of text entities comprising part-of-speech (POS)-based extraction, and machine learning (ML)-based extraction, wherein performing the ML-based extraction comprises Generating recommendations based on historical data using a classification-based ML model Causing at least one of machine learning (ML)-based extraction including The system according to claim 12. **Claim 15** Identifying text data from an input file, wherein the identified text data includes a plurality of text entities; Receiving text input from a user to identify related text entities from the plurality of text entities; Automatically generating a search pattern corresponding to the text input; Determining a pattern associated with each of the plurality of text entities; Associating the search pattern corresponding to the text input with the pattern associated with the plurality of text entities; Identifying one or more matching patterns from the patterns associated with the plurality of text entities based on the association; Extracting related text entities corresponding to the one or more matching patterns from the plurality of text entities A non-transitory computer-readable storage medium storing a set of computer-executable instructions that cause a computer including one or more processors to perform one or more steps including. **Claim 16** The set of computer-executable instructions causes the computer including the one or more processors to Domain-based extraction, wherein performing the domain-based extraction comprises Determining a domain-based name associated with each of the plurality of text entities, the domain-based name being determined based on associating each of the plurality of text entities with a list of domain-based names stored in a data repository; determining Domain-based extraction including Location-based extraction, where performing the location-based extraction Identifies the type associated with the input file, Based on the type associated with the input file, determines the location of one or more relevant text entities in the input file, and Based on the location, extracts the one or more relevant text entities from the input file, where the location is predetermined based on a template corresponding to the type associated with the input file stored in a data repository, extraction Location-based extraction including Part-of-speech (POS)-based extraction, where performing the part-of-speech (POS)-based extraction Identifies the POS associated with the plurality of text entities, where the POS is one of a noun, pronoun, verb, adverb, adjective, conjunction, preposition, and interjection, identification, and Extracts one or more text entities where the POS is identified as a noun, Identifies one or more unique text entities where the POS is not identified, Determines the entity type associated with each of the one or more unique text entities, where the entity type is one of a value entity and a pattern entity, each value entity has an associated value, and each pattern entity has an associated pattern, determination For each value entity, Automatically identifies one or more text entities having the same associated value as the associated value of the value entity, For each pattern entity, Automatically generates a search pattern corresponding to the pattern associated with the pattern entity, Associating the search pattern corresponding to the pattern entity with a pattern associated with the plurality of text entities; Determining one or more matching patterns from the patterns associated with the plurality of text entities based on the association, and Extracting matching text entities corresponding to the one or more matching patterns from the plurality of text entities including part-of-speech (POS)-based extraction, and machine learning (ML)-based extraction, wherein performing the ML-based extraction comprises generating recommendations based on historical data using a classification-based ML model causing at least one of the machine learning (ML)-based extractions to be performed; The non-transitory computer-readable storage medium according to claim 15.

Citation Information

Patent Citations

  • Information processing apparatus and information processing program

    JP2014109810A

  • Information processing device and program

    JP2019185631A

  • Data capture from images of documents with fixed structure

    US20150278593A1

  • System and method for data extraction and searching

    US20190286898A1