Information expression structure analysis device and information expression structure analysis method

The information expression structure analysis device and method enhance document analysis by utilizing control characters and HTML tags, addressing inefficiencies in existing systems to improve the extraction of information from unstructured documents.

JP7795869B2Active Publication Date: 2026-01-08HITACHI LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021065806
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-04-08
Publication Date
2026-01-08
Estimated Expiration
2041-04-08

AI Technical Summary

Technical Problem

Existing systems for document analysis, such as those using OCR and two-dimensional statistical analysis, fail to utilize control characters, HTML tags, and other structural information, leading to inefficient extraction of desired information from unstructured documents.

Method used

An information expression structure analysis device and method that utilize an information processing device to identify and extract information from unstructured documents by considering text data, structural information, and control characters, using information expression grammars and templates to generate patterns for efficient extraction.

Benefits of technology

Enables efficient extraction of desired information from unstructured documents by leveraging various forms of document information, including control characters and HTML tags, improving the accuracy and completeness of document analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007795869000001
    Figure 0007795869000001
  • Figure 0007795869000002
    Figure 0007795869000002
  • Figure 0007795869000003
    Figure 0007795869000003
Patent Text Reader

Abstract

To efficiently extract target information from a non-fixed form document.SOLUTION: An information expression structure analysis apparatus stores an information expression template which is a template used in generating an information expression pattern which is a program code for realizing a function of extracting an extraction target for each combination of an information expression grammar and a support information type which is a classification of support information used when extracting an extraction target, identifies the information expression template used in generating the information expression pattern used in extracting an extraction target from a non-fixed form document based on the information expression grammar and the support information type identified for an information expression, and generates the information expression pattern by applying the extraction target and basis information which is a basis for extracting the extraction target from the information expression to the identified information expression template.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information expression structure analysis device and an information expression structure analysis method. [Background technology]

[0002] Patent Document 1 describes a method for scanning printed materials and typed materials using an image scanner. Optical character recognition (OCR) is performed on the image read by A system is described that aims to enable accurate and efficient recognition of documents when extracting text using GraphQL technology. The system recognizes the layout structure of a document (columns, author, title, footnotes, etc.) by grammatically analyzing the visual structure of the document using two-dimensional adaptation of statistical analysis algorithms, and interprets the structural components of the document. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Special Publication No. 2009-500755 Summary of the Invention [Problem to be solved by the invention]

[0004] The system described in Patent Document 1 recognizes the layout structure of a document by grammatically analyzing the visual structure of the document and interprets the structural components of the document. However, the system described in this document does not utilize information other than the two-dimensional arrangement of text and text within the text, such as control characters contained in the document, such as tables, spaces, tabs, and HTML tags (HyperText Markup Language), information written outside the document, such as headers and footers, and invisible information in the document (information that does not appear on the surface of the document), which are useful clues for interpreting the structure of the document. As a result, it is not always possible to efficiently extract desired information from unstructured documents.

[0005] The present invention has been made in consideration of the above background, and aims to provide an information expression structure analysis device and an information expression structure analysis method that enable efficient extraction of desired information from non-standard documents. [Means for solving the problem]

[0006] One aspect of the present invention for solving the above problems is an information expression structure analysis device, which is configured using an information processing device having a processor and a storage device, The storage device includes: storing information expressions, which are the manner of expressing information in an unstructured document including text data and structural information; extraction targets, which are information that a user intends to extract from the unstructured document; and grounds information, which are information used to determine whether or not to extract the extraction targets from the information expressions; The processor: extracting text data related to the extraction target from the information expression and extracting the structural information related to the extracted text data; The processor: The number of extracted text data, the basis information whether the basis information is within a specified range; and whether the basis information is information or a group of information that cannot be compared in magnitude with the information extraction target in the information expression. Based on the above, the type of information expression grammar that describes the information expression is Is it one of the following (1) to (9)? Identify the (1) The number of extracted text data is one, and the basis information has not been acquired. (2) The number of extracted text data is one, and the basis information has been acquired, and the basis information is a range specification. (3) The number of extracted text data is one, and the basis information has been acquired, and the basis information is not a range specification, and the basis information is information or a group of information that cannot be compared in size with the extraction target in information expression. (4) The number of extracted text data is one, and the basis information has been acquired, and the basis information is not a range specification, and the basis information is not information or a group of information that cannot be compared in size with the information extraction target in information expression. (5) The number of extracted text data is two or more, and the basis information has not been acquired, and multiple extraction targets are information or a group of information that can be compared in size in information expression. (6) The number of extracted text data is two or more, and the basis information has not been acquired, and multiple extraction targets are not information or a group of information that can be compared in size in information expression. (7) The number of extracted text data is two or more, and the basis information has been acquired, and the basis information is a range specification. (8) The number of extracted text data is two or more, the basis information is obtained, the basis information is not a range specification, and the multiple extraction targets are information or information groups that can be compared in terms of information expression. (9) The number of extracted text data is two or more, the basis information is obtained, the basis information is not a range specification, and the multiple extraction targets are not information or information groups that can be compared in terms of information expression. The processor: Based on the identified type of information expression grammar and type of structural information, a support information type, which is a classification based on the structure of the information expression, is: It is determined whether the information expression grammar is one of the following (a) to (f): (a) the information expression grammar is the information expression grammar of (1) above and matches a regular expression; (b) the information expression grammar is the information expression grammar of (1) above and the loaded dictionary words and the extraction target match; or (c) the information expression grammar is the information expression grammar of (1) above and the loaded dictionary words and the extraction target do not match, and meta-information is extracted from the unstructured document, and the basis information exists in the extracted meta-information. (d) When the information expression grammar is the information expression grammar of (1), the loaded dictionary word does not match the extraction target, meta-information is extracted from the unstructured document, the basis information is not present in the extracted meta-information, and the unstructured document has been formatted with predetermined control characters embedded, or when the information expression grammar is not any of (1), (4), (8), or (9), the extraction target and the basis information cannot be compared in terms of magnitude, and the unstructured document has been formatted with predetermined control characters embedded. (e) When the information expression grammar is any of (4), (8), or (9), or when the information expression grammar is not any of (1), (4), (8), or (9), and the extraction target and the basis information can be compared in terms of magnitude. (f) The information expression grammar is not any of (1), (4), (8), or (9), the extraction target and the basis information are not comparable in magnitude, and the unstructured document is not a document formatted with a predetermined control character embedded therein.storing an information expression template table including, for each of the information expression grammars, information expression templates that are templates of information expression patterns that are program codes that realize a function of acquiring an extraction target from the unstructured document, the information expression templates being described in association with the support information type; The processor: obtaining the information expression template corresponding to the combination of the identified information expression grammar and the identified support information type from the information expression template; The processor replaces the symbol string of the information expression template with the extraction target and the basis information. By doing so, the information expression pattern is generated.

[0007] Other problems and solutions disclosed in the present application will be made clear in the detailed description and drawings. [Effects of the Invention]

[0008] According to the present invention, it is possible to efficiently extract desired information from unstructured documents. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating a schematic configuration of a document information management system. [Figure 2] 1 is an example of an information processing apparatus that constitutes a document information management system. [Figure 3] This is an example of an unconventional document. [Figure 4] This is an example of an information expression pattern. [Figure 5] 10 is an example of an information expression template. [Figure 6] 10 is an example of an information expression template table. [Figure 7] 10 is a flowchart illustrating an information expression pattern generation process. [Figure 8] 10 is a flowchart illustrating an information expression grammar identification process. [Figure 9] 10 is a flowchart illustrating a support information type identification process. [Figure 10]FIG. 10 is a diagram illustrating a schematic configuration of a document information management system according to a second embodiment. [Figure 11] 10 is a flowchart illustrating a specific support information acquisition process. [Figure 12] 10 is an example of a specific support information acquisition screen. [Figure 13] 10 is an example of a specific support information acquisition screen. [Figure 14] 10 is an example of a specific support information acquisition screen. [Figure 15] 10 is an example of a specific support information acquisition screen. [Figure 16] 10 is an example of a specific support information acquisition screen. [Figure 17] 10 is an example of a specific support information acquisition screen. [Figure 18] 10 is an example of a specific support information acquisition screen. [Figure 19] FIG. 10 is a diagram illustrating a schematic configuration of a document information management system according to a third embodiment. [Figure 20] 10 is a flowchart illustrating an information expression pattern verification process. [Figure 21] 10 is an example of an information expression pattern verification result display screen. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments will be described with reference to the drawings. Note that the following description and drawings are merely examples for explaining the present invention, and some details have been omitted or simplified as appropriate for clarity of explanation. The present invention can be implemented in various other forms. Unless otherwise specified, each component may be singular or plural.

[0011] In the following explanation, identical or similar components may be assigned the same reference numerals, and redundant explanations may be omitted. Furthermore, in the following explanation, the letter "S" before a reference numeral indicates a processing step. Furthermore, in the following explanation, various types of information may be described using expressions such as "information," "data," and "table," but the various types of information may also be handled using data structures other than those illustrated.

[0012] [First embodiment] FIG. 1 shows a schematic configuration of an information processing system (hereinafter referred to as "document information management system 1") described as a first embodiment. As shown in the figure, the document information management system 1 includes an unstructured document management device 2, a user device 3, and an information expression structure analysis device 100. All of these are configured using information processing devices (computers), and are connected to each other via a communication network 5 in a state where they can communicate with each other bidirectionally. The communication network 5 may be, for example, a LAN (Local Area Network), a WAN (Wide Area Network), an Internet These include the Internet, dedicated lines, and various public communication networks.

[0013] The non-standard document management device 2 extracts information that the user wants to extract (hereinafter referred to as "extracted information") from non-standard documents (documents whose formats vary depending on the issuer, such as reports, statements, financial statements, and various registration forms (e.g., Rich Text Format)). Hereinafter referred to as "non-standard documents"), and provides the extracted information to the user via the user device 2.

[0014] An unstructured document is a collection of text data (hereinafter referred to as "text") of words and sentences, structural information (tables, spaces, tabs, information written outside the document, information invisible on the document (regular expressions, dictionary matches, meta information, HTML (HyperText Markup Language) tags, etc.), and so on. Control characters), etc. Hereinafter referred to as "structural information". Hereinafter, the manner of expression of information in an unstructured document (expression of information by text or structural information) contained in an unstructured document will be collectively referred to as "information expression". Information expression can be, for example, document data of a specific data type handled by application software such as word processing software, data describing a web page, or data generated by optical character recognition (OCR) technology. The image data is extracted from the image data (image data acquired by an image scanner).

[0015] As shown in the figure, the unstructured document management device 2 has the functions of an unstructured document management unit 21, an information extraction unit 22, an extracted information management unit 23, and an extracted information provision unit 24.

[0016] Of these, the non-standard document management unit 21 acquires non-standard documents through user input via the user device 3 or through provision from other information processing devices via the communication network 5, and manages the acquired non-standard documents.

[0017] The information extraction unit 22 acquires (extracts) extracted information from unstructured documents. The information extraction unit 22 acquires extracted information from unstructured documents by executing (performing pattern matching of information expression) program code (or pseudo code) (hereinafter referred to as "information expression pattern") that realizes the function of acquiring extracted information from unstructured documents. This information expression pattern is generated by the information expression structure analysis device 100. The information expression pattern can also be edited by the user via the user device 3.

[0018] The extraction information management unit 23 manages the extraction information acquired by the information extraction unit 22. The extraction information provision unit 24 provides the extraction information managed by the extraction information management unit 23 to the user device 3.

[0019] The user device 3 has the functions of a various setting unit 31 and an extracted information utilization unit 32. The various setting unit 31 performs various settings required for the information expression structure analysis device 100 to generate and edit information expression patterns. The extracted information utilization unit 32 requests extracted information requested by the user from the unstructured document management device 2, receives the extracted information sent from the unstructured document management device 2, and provides it to the user.

[0020] The information expression structure analysis device 100 generates an information expression pattern and provides it to the unstructured document management device 2. As shown in the figure, the information expression structure analysis device 100 has the functions of a memory unit 110, an information expression structure analysis unit 120, and an information expression pattern generation unit 130.

[0021] As shown in the figure, the storage unit 110 stores extraction target information 101, a basis information group 102, an information expression group 111, an information expression template group 112, an information expression template table 113, an information expression pattern group 114, and various dictionaries 115.

[0022] The extraction target information 101 includes one or more extraction targets that the user intends to extract from an unstructured document. The extraction target information 101 is set by the user via the user device 2, for example.

[0023] The basis information group 102 includes one or more pieces of basis information that serve as a basis for extracting an extraction target from an information expression. The basis information group 102 is set by the user via the user device 2, for example.

[0024] The information expression group 111 includes one or more information expressions extracted from unstructured documents. The information expression group 111 is set by a user, for example, via the user device 2. For example, when attempting to generate an information expression pattern to be used for extracting extracted information from a large number of unstructured documents, the user registers the information expressions extracted from those unstructured documents as the information expression group 111 in the information expression structure analysis device 100.

[0025] The information expression template group 112 includes one or more program codes (hereinafter referred to as "information expression templates") that are templates of information expression patterns. The information expression templates will be described in detail later.

[0026] The information expression template table 113 is referred to by the information expression structure analysis unit 120 when selecting an information expression template to be used for generating an information expression pattern.

[0027] The information expression pattern group 114 includes one or more information expression patterns generated by the information expression pattern generating unit 130 .

[0028] The various dictionaries 115 include various dictionaries (word dictionaries, regular expression dictionaries, etc.) used by the information expression structure analysis unit 120 and the information expression pattern generation unit 130.

[0029] The information expression structure analysis unit 120 identifies information (information expression grammar and support information type, described later) to be used to search for an information expression template from the information expression template table 113 based on the information expressions (text, structural information) of the information expression group 111.

[0030] As shown in the figure, the information expression structure analysis unit 120 has the functions of a text information extraction unit 121, a structural information extraction unit 122, an information expression grammar identification unit 123, and a support information type identification unit .

[0031] Of these, the text information extraction unit 121 extracts text from the information expression. The structural information extraction unit 122 extracts structural information from the information expression. The information expression grammar identification unit 123 identifies a grammar for describing the information expression (hereinafter referred to as "information expression grammar") based on the text and structural information extracted from the information expression. The support information type identification unit 124 identifies a support information type (described below) corresponding to the information expression from the information expression template table 113 based on the text and structural information extracted from the information expression.

[0032] The information expression pattern generation unit 130 shown in Figure 1 searches for an information expression template from the information expression template table 113 based on the information expression grammar and support information type identified by the information expression structure analysis unit 120, and generates an information expression pattern by applying a specific extraction target (text, etc.) and basis information (composition information, etc.) to the searched information expression template.

[0033] As shown in the figure, the information expression pattern generation unit 130 has the functions of an information expression template search unit 131 and an information expression component replacement unit 132 .

[0034] Of these, the information expression template search unit 131 searches the information expression template table 113 for an information expression template that corresponds to the combination of the information expression grammar and the type of support information identified by the information expression structure analysis unit 120 .

[0035] In addition, the information expression component replacement unit 132 generates an information expression template by applying (substituting) a specific extraction target (text, etc.) and basis information (composition information, etc.) to the information expression template searched by the information expression template search unit 131.

[0036] 2 shows an example of the hardware configuration of information processing devices (unstructured document management device 2, user device 3, and information expression structure analysis device 100) that make up the document information management system 1. The illustrated information processing device 10 includes a processor 11, a main memory device 12, an auxiliary memory device 13, an input device 14, an output device 15, and a communication device 16. The information processing device 10 is, for example, a personal computer, an office computer, a server device, a smartphone, a tablet, or the like.

[0037] All or part of the information processing device 10 may be realized using virtual information processing resources provided using virtualization technology, process space separation technology, etc., such as a virtual server provided by a cloud system. Also, all or part of the functions provided by the information processing device 10 may be realized by a service provided by the cloud system via an API (Application Programming Interface), etc. Furthermore, all or part of the functions provided by the information processing device 10 may be realized using, for example, SaaS (Software as a Service), PaaS (Platform as a Service), IaaS (Infrastructure as a Service), etc.

[0038] For example, at least two or more of the unstructured document management device 2, the user device 3, and the information expression structure analysis device 100 may be realized by the same information processing device 10 (common hardware).

[0039] The processor 11 shown in the figure is, for example, a CPU (Central Processing Unit), an MPU (Microprocessor Unit), (Micro Processing Unit), GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit), It is constructed using AI (Artificial Intelligence) chips, etc.

[0040] The main memory device 12 is a device for storing programs and data, and is, for example, a ROM (Read Only Memory). These include non-volatile memory (NVRAM (Non Volatile RAM)), RAM (Random Access Memory), and non-volatile memory (NVRAM).

[0041] The auxiliary storage device 13 is, for example, an SSD (Solid State Drive), a hard disk drive, The auxiliary storage device 13 may be a hard disk, an optical storage device (e.g., a CD (Compact Disc), a DVD (Digital Versatile Disc)), a storage system, a reading / writing device for a recording medium such as an IC card, an SD card, or an optical recording medium, or a storage area of ​​a cloud server. Programs and data can be read into the auxiliary storage device 13 via a recording medium reading device or a communication device 16. The programs and data stored (memorized) in the auxiliary storage device 13 are read into the main storage device 12 as needed.

[0042] The input device 14 is an interface that accepts input from the outside, and is, for example, a keyboard, a mouse, a touch panel, a card reader, a pen-input tablet, a voice input device, or the like.

[0043] The output device 15 is an interface that outputs various information such as the progress of processing and the results of processing. The output device 15 is, for example, a display device (liquid crystal monitor, LCD (Liquid Crystal Display), graphic card, etc.) that visualizes the various information described above, a device that converts the various information described above into audio (audio output device (speaker, etc.)), or a device that converts the various information described above into text (printer, etc.). Note that, for example, the information processing device 10 may be configured to input and output information to and from other devices via the communication device 16.

[0044] The input device 14 and the output device 15 constitute a user interface that realizes interactive processing with the user (receiving information, presenting information, etc.).

[0045] The communication device 16 is a device that realizes communication with other devices. The communication device 16 is a wired or wireless communication interface that realizes communication with other devices via the communication network 5, and is, for example, a NIC (Network Interface Card), a wireless communication module, a USB module, or the like.

[0046] The information processing device 10 may be equipped with, for example, an operating system, a file system, a DBMS (DataBase Management System) (relational database, NoSQL, etc.), a KVS (Key-Value Store), etc.

[0047] The functions of the non-standard document management device 2, the user device 3, and the information expression structure analysis device 100 are realized by the processor 11 of each of the information processing devices 10 reading and executing a program stored in the main memory device 12 of each of the information processing devices 10, or by the hardware (FPGA, ASIC, AI chip, etc.) of each of the information processing devices 10 itself.

[0048] The unstructured document management device 2, the user device 3, and the information expression structure analysis device 100 store the various pieces of information (data) described above, for example, as tables in a database or files managed by a file system.

[0049] An example of an unstructured document is shown in Figure 3. The illustrated unstructured document 300 is a document that an organization such as a company submits when applying for a loan to a financial institution. The unstructured document 300 contains text and various structural information (control characters such as tables, spaces, tabs, and HTML tags, information written outside the document, information invisible on the document, etc.) as information representation.

[0050] As shown in the figure, the illustrated unstructured document 300 has the following columns: a header 310, a body 320, and a footer 330. The header 310 contains an application date 301. The body 320 contains a company registration date 321 and a financial statement 322 showing the company's financial status. The footer 330 contains a page number 331.

[0051] For example, company registration date 321 written in body 320 includes the text "Registration" and "07 / 16 / 2007" as information expressions, and structural information that these are written on the same line. For example, by treating this information expression as a pattern and generating an information expression pattern (program code), it becomes possible to automatically obtain company registration dates from a variety of unstructured documents by pattern matching.

[0052] 4 shows an example of an information expression pattern. The illustrated information expression pattern 400 implements a function that uses a date (date) 410 included in a document as input and determines whether the result of a function "same_line_word" 420, which retrieves words that exist on the same line as the date (date) 410, contains the word "Registration." The illustrated information expression pattern 400 returns "TRUE" if the execution result of the function "same_line_word" 420 contains the word "Registration," and returns "FALSE" if the word is not included.

[0053] Fig. 5 shows an example of an information expression template used to generate the information expression pattern 400 shown in Fig. 4. The example information expression template 500 includes columns for type "Type" 510, description "Description" 520, and template code "Template Code" 530.

[0054] The type "Type" 510 contains elements of the information expression grammar for the information expression targeted by the information expression template 500 (in this example, C (information extraction target), [Position], and [Word]). The initial letter of the current element and the symbol string in square brackets (in this example, "C[P][W]") are Note that the [Position] and [Word] in square brackets are subject to replacement when converting from an information expression template to an information expression pattern.

[0055] The description 520 explains the logic of the information expression template in natural language. In this example, the description "Description" 520 describes information indicating that "C (information extraction target) is in a predetermined positional relationship (Position) with a word or a group of words."

[0056] The template code "Template Code" 530 contains a program that represents the logic of information expression. The template of the ramcode is written. By replacing the part in square brackets in "Template Code" 530 with specific content, an information expression pattern is generated. For example, if the information expression is the company registration date 321 of the unstructured document 300 illustrated in FIG. 3, the [Word] in the template code "Template Code" 530 is replaced with the word "Registration." Replace [Position] with "in same_line" and C with "date". This generates the information expression pattern 400 shown in Fig. 4. Note that function(Position) is a function that returns a function that acquires a word group based on the argument "Position." In the case of Fig. 4, the argument "Position" is replaced with "in same_line," which means the same line, so function(Position) returns the function "same_line_word" 420 shown in Fig. 4.

[0057] 6 shows an example of the information expression template table 113. The illustrated information expression template table 113 is made up of multiple entries (records) each having the following items: row number 1131, grammar expression element 1132, information expression grammar 1133, and support information type 1134. One entry in the information expression template table 113 corresponds to one information expression grammar.

[0058] Of the above items, the row number 1131 stores an identifier (row number) assigned to each entry in the information expression template table 113.

[0059] The grammar expression element 1132 stores the above-mentioned grammar expression element of the information expression grammar.

[0060] The information expression grammar 1133 stores the content of the information expression grammar expressed in a predetermined notation method (for example, a notation method conforming to the notation method of natural language grammar).

[0061] Support information type 1134 is set with a support information type, which is a classification based on the structure of information expression of support information, which is information used when extracting an extraction target from an information expression. In addition to distinctions based on information expression grammar, support information types also provide a more detailed classification of information expression templates used to generate information expression patterns. In this example, the support information types are exemplified as "regular expression," "dictionary match," "meta information (page number, etc.)," ​​"HTML structure," "set," and "structure." By classifying information expression templates according to the exemplified support information types, it is possible to comprehensively handle non-standard documents of various formats.

[0062] In the figure, nine information expression grammars are illustrated, each distinguished by line numbers #1 to #9.

[0063] Among these, the line number "#1" is Information Expression Grammar 1133 "C is the [Same] as [Words]." The content of the grammar expression element 1132 of the line is "C[S][W]" which indicates the grammar expression element of the information expression grammar. In this way, the extraction target and the predetermined information (word) or information group (word group) have the same meaning [Same]. By using information expression grammar that expresses the semantic relationship between information or groups of information as one of the classifications, This makes it possible to efficiently identify appropriate information expression templates.

[0064] Of the support information types 1134 for the row, "regular expression," "dictionary match," "meta information," and "HTML structure" store information expression template identifiers "Template1" to "Template4," respectively.

[0065] Information expression template "Template1" stored in the support information type "Regular expression" contains the program code that determines whether the expression in C (the object of information extraction) matches the regular expression. For example, in the case of the unstructured document 300 in FIG. 1, the above program code extracts "numerical value / " for the date "07 / 16 / 2007" which is C (information extraction target). Determines whether three numbers, such as "number / number," match a regular expression separated by slashes.

[0066] The information expression template "Template2" stored in the support information type "Dictionary Match" " indicates whether C (information extraction target) matches a word included in a dictionary set by the user, etc. For example, in the case of the unstructured document 300 in FIG. 1, the program code is written as follows: It is determined whether the word matches a word contained in a dictionary of company names.

[0067] Information expression template "Template3" stored in the support information type "meta information" In this case, C (information extraction target) is not expressed as a string on the document, for example, the total number of pages or The template of the program code is written to determine whether or not the character strings contained in the creator, creation date, etc. of the child document match. For example, in the case of the unstructured document 300 in FIG. 1, C (information extraction target) is the date "07 / 16 / 2007" that appears on the last page of multiple pages. If so, the program code determines whether the total number of pages in the document matches the number of pages on which the date "07 / 16 / 2007" exists.

[0068] The information expression template "Template4" stored in the support information type "HTML structure" shows that the non-standard document is an HTML-like document with embedded control characters. If so, whether the control character for displaying C (information extraction target) matches a specific control character For example, when the unstructured document 300 in FIG. 1 is written in HTML, C (information extraction target) is written in the header 310. If the application date exists, "09 / 22 / 2010", the above program code will <header>Determine whether the date tagged as a control character matches C (information extraction target). do.

[0069] The information expression grammar 1133 with line number "#2" "C is the [Position] of [Ranges]." means that "the position [Position] of C (information extraction target) belongs to a region or a group of regions [Ranges]." The grammar expression element 1132 of the line stores "C[P][R]" which indicates the grammar expression element of the information expression grammar. In this way, the relationship between the position of the extraction target and the region or group of regions is By using the information expression grammar representing the above as one of the classifications, it becomes possible to efficiently identify appropriate information expression templates.

[0070] In the support information type 1134 of the row, "HTML structure", "set", and "structure" store the information expression template identifiers "Template5" to "Template7", respectively.

[0071] The information presentation template "Template5" stored in the support information type "HTML structure" assumes a document like HTML that has been formatted with embedded control characters, and The information extraction target exists in a specific positional relationship [Position] with the ranges [Ranges] expressed in HTML. For example, if the unstructured document 300 in FIG. 1 is a document written in HTML, the program code determines whether or not C (information extraction target) that is the date "07 / 16 / 2007" below the header 310 is HTML <header>It is determined whether it exists after the end position of the tag.

[0072] The information expression template "Template6" stored in the support information type "Collection" contains , determine whether C (information extraction target) exists in a specific positional relationship [Position] with the range group [Ranges]. For example, C (information extraction target) is the date "07 / 16 / 2007" near the center of the body 320 of the unstructured document 300 in FIG. In this case, the program code determines, for example, whether the target date is between the upper 20% and lower 20% regions of the body 320.

[0073] The information expression template "Template7" stored in the support information type "Structure" contains , C (information extraction target) exists in a specific positional relationship [Position] with one of the ranges [Ranges]. For example, if C (information extraction target) is the application date "09 / 22 / 2010" located above the unstructured document 300 in FIG. 1, the program code determines whether the target date is "09 / 22 / 2010" or not. Determine whether it is in the upper 10% range of

[0074] The information expression grammar 1133 "C is the [Position] of [Words]." with line number "#3" is The content is "The position [Position] of C (information extraction target) has a predetermined positional relationship with the predetermined information (word) or information group (word group) [Words]." The grammar expression element 1132 of the line stores "C[P][W]" which indicates the grammar expression element of the information expression grammar. By using an information expression grammar that expresses the positional relationship between the position of the target and predetermined information or information group as one of the classifications, it becomes possible to efficiently identify an appropriate information expression template.

[0075] The "HTML structure", "set", and "structure" of the support information type 1134 of the row in question store the information expression template identifiers "Template8" to "Template10", respectively. In addition, these information expression templates are information expression templates "Template5" to "Template7" with line number "#2", where ranges are replaced with words. This is what happened.

[0076] For example, in the information expression template "Template8" stored in the support information type "HTML structure," C (information extraction target) is the year "2007" of the non-standard document 300 in FIG. In this case, a template of program code is written that determines whether or not the word "Year" exists in a tag that exists in the same column.

[0077] The information expression grammar 1133 "C is the [Relation] of [Words]" with the line number "#4" has the content that "C (information extraction target) has a predetermined relationship [Relation] with a predetermined information (word) or information group (word group) [Words]." The grammar expression element 1132 of the line stores "C[R][W]" which indicates the grammar expression element of the information expression grammar. In this way, the position of the extraction target By using an information expression grammar that expresses the relationship between the information and a predetermined piece of information or a group of information as one of the classifications, it becomes possible to efficiently identify an appropriate information expression template.

[0078] The identifier of the information expression template, "Template11", is stored in the "set" of the support information type 1134 of the relevant row. It is determined whether or not a specific relationship [Relation] exists with a group of words [Words]. Here, the relationship [Relation] refers to a relationship specified by a comparison operation such as "equal," "greater than," "smaller than," "maximum," or "minimum." For example, if C (information extraction target) is the latest year "2009" of the unstructured document 300 in FIG. 1, the information expression template may contain the word "Year." A template of program code is described that determines whether "2009" is the latest (largest value) of the words "2007," "2008," and "2009" that exist in the same column.

[0079] The information expression grammars with line numbers "#5" to "#9" all contain multiple Cs (information extraction targets). This is the case.

[0080] The information expression grammar 1133 with the line number "#5" "C is the [Position] of [Words(C)]." means that "the position [Position] of the first C (information extraction target) is in a predetermined positional relationship with the information (word) or information group (word group) [Words] that is the second C (information extraction target)." The grammar expression element 1132 of the line contains "C[P][C]" which indicates the grammar expression element of the information expression grammar. In this way, by using the information expression grammar that represents the positional relationship between the first and second extraction targets as one of the classifications, it becomes possible to efficiently identify an appropriate information expression template.

[0081] In the support information type 1134 of the row, "HTML structure", "set", and "structure" store the information expression template identifiers "Template12" to "Template14", respectively.

[0082] The information expression template "Template12" stored in the support information type "HTML structure" describes a template for program code that determines whether a first C (information extraction target) is in a specific positional relationship [Position] with a second C (information extraction target) to be combined, assuming a document formatted with embedded control characters such as HTML. For example, if the unstructured document 300 in FIG. 1 is written in HTML and the first C (information extraction target) is "2007" and the second C (information extraction target) is "$100,000," the program code extracts sales information for each fiscal year by determining whether the first and second C (information extraction targets) are in a specific positional relationship [Position]. It determines whether the HTML tags are adjacent or not.

[0083] The information expression template "Template13" stored in the support information type "Set" has a specific positional relationship between the first C (information extraction target) and a group of words [Words] containing the second C (information extraction target). For example, when extracting information to calculate the average value of profit (Net Profit) from the unstructured document 300 in FIG. 1, for example, the first C (information extraction target) is "Net Profit", the second C is "Net Profit", and the third C is "Position". If C (information extraction target) is either "$10,000", "$30,000", or "$30,000", the above program code checks whether the first and second C (information extraction target) exist in the same column. Make a judgment.

[0084] The information expression grammar 1133 with line number "#6," "C & C is the [Position] of [Ranges].", states that "The positions [Position] of the first C (information extraction target) and the second C (information extraction target) both belong to a predetermined region or region group [Ranges]." The grammar expression element 1132 of that line stores "CC[R][W]," which indicates the grammar expression element of that information expression grammar. In this way, by using the information expression grammar that expresses the relationship between the respective positions of the first extraction target and the second extraction target and a predetermined region or regions as one of the classifications, it becomes possible to efficiently identify an appropriate information expression template.

[0085] In the support information type 1134 of the row, "HTML structure", "set", and "structure" store the information expression template identifiers "Template15" to "Template17", respectively.

[0086] The information expression template "Template15" stored in the support information type "HTML structure" assumes that the document is formatted with embedded control characters like HTML, and A template of program code is described that determines whether the first C (information extraction target) and the second C (information extraction target) are in a specific positional relationship [Position] with one of the ranges [Ranges]. For example, if the unstructured document 300 in FIG. 1 is written in HTML, and the first C (information extraction target) is "2007" and the second C (information extraction target) is "$100,000", the above program code will, for example, extract the sales for each fiscal year by determining whether the first and second C (information extraction targets) are in a specific positional relationship [Position] with one of the ranges [Ranges]. It also checks whether they are contained in the same HTML tag.

[0087] The information expression template "Template16" stored in the support information type "Set" describes a template of program code that determines whether the first C (information extraction target) and the second C (information extraction target) are in a specific positional relationship [Position] with the area group [Ranges]. For example, in the unstructured document 300 of FIG. 1, the first C (information extraction target) is "2007" and the second C is "2008". If C (information extraction target) of 2 is "$100,000", the above program code will be, for example, The first and second C (information extraction target) are both areas in which the column names of the unstructured document 300 are written. It is also determined whether or not the area is included in the first column to which serial numbers are assigned.

[0088] The information expression template "Template17" stored in the support information type "structure" describes a template of program code that determines whether the first C (information extraction target) and the second C (information extraction target) are in a specific positional relationship [Position] with one of the ranges [Ranges]. For example, in the unstructured document 300 of FIG. 1, if the first C (information extraction target) is "2007 ", if the second C (information extraction target) is "$100,000", the above program code will For example, if the first and second C (information extraction target) both contain the column names of the unstructured document 300, It is determined whether the area is below the area where the

[0089] Line number "#7" in Information Expression Grammar 1133: "C & C is the [Position] of [Words]." The definition is "The positions of the first C (information extraction target) and the second C (information extraction target) both have a specific positional relationship with the predetermined information (word) or information group (word group) [Words]." The grammar expression element 1132 of the relevant line stores "CC[P][W]" which indicates the grammar expression element of the relevant information expression grammar. In this way, by using the information expression grammar which indicates the positional relationship between the respective positions of the first extraction target and the second extraction target and the predetermined information or information group as one of the classifications, it becomes possible to efficiently identify an appropriate information expression template.

[0090] The "HTML structure," "set," and "structure" fields in the support information type 1134 for the row in question store the information presentation template identifiers "Template18" to "Template20," respectively. These information presentation templates "Template18" to "Template20" are obtained by replacing the range group [Ranges] with the word group [Words] in the information presentation templates "Template15" to "Template17" in row number "#6," respectively.

[0091] For example, in the information expression template "Template18" stored in the support information type "HTML structure", the first C (information extraction target) and the second C (information extraction target) are the word group [Words]. For example, if the unstructured document 300 in FIG. 1 is written in HTML, and the first C (information extraction target) is "2007" and the second C (information extraction target) is "$100,000," the program code will determine whether the first and second C (information extraction target) are both consecutive numbers. Determine whether it exists in the same column as the number "1".

[0092] The information expression grammar 1133 "C is the [Relation] of C." in line number "#8" is The first C (information extraction target) and the second C (information extraction target) have a predetermined relationship [Relation]." The grammar expression element 1132 of the relevant line stores "C[R]C" which indicates the grammar expression element of the relevant information expression grammar. Here, the relationship [Relation] refers to a relationship specified by a comparison operation such as "equal to," "greater than," "less than," "maximum," or "minimum." In this way, by using the information expression grammar that expresses the relationship between the first extraction target and the second extraction target as one of the classifications, it becomes possible to efficiently identify an appropriate information expression template.

[0093] The identifier of the information expression template "Template21" is stored in "Set" of the support information type 1134 of the row. This information expression template "Template21" describes a template of program code that determines whether or not the first C (information extraction target) and the second C (information extraction target) have a specific relationship [Relation]. For example, when extracting a combination of sales and net profit from the unstructured document 300 in FIG. 1, 2007 Regarding the sales and profits for the fiscal year, the first C (information extraction target) is "$100,000" and the second C (information extraction If the target C is "$100,000", the above program code determines, for example, that a larger numerical value C (information extraction target) is sales, and a smaller numerical value C (information extraction target) is profit.

[0094] Line number "#9" in Information Expression Grammar 1133: "C & C is the [Relation] of [Words]." The content of the line is "The first C (information extraction target) and the second C (information extraction target) have a predetermined relationship [Relation] with the predetermined information (word) or information group (word group) [Words]." The grammar expression element 1132 stores "CC[R][W]" which indicates the grammar expression element of the information expression grammar. Here, the relation [Relation] can be, for example, "equivalent," "greater than," or "less than." This refers to a relationship specified by a comparison operation such as "smallest," "largest," or "smallest." By classifying the information expression grammar that expresses the relationship between the specified information or information groups of the first extraction target and the second extraction target, it becomes possible to efficiently specify an appropriate information expression template.

[0095] The identifier of the information expression template "Template22" is stored in the "Set" of the support information type 1134 of the row. This information expression template "Template22" describes a template of program code that determines whether or not the first C (information extraction target) and the second C (information extraction target) have a specific relationship [Relation] with a specific word group [Words]. For example, when a combination of sales and net profit is extracted from the unstructured document 300 in FIG. When extracting sales and profits for fiscal year 2007, if the first C (information extraction target) is "$100,000" and the second C (information extraction target) is "$100,000," in order to separate it from the serial number and fiscal year figures, the above program code, for example, compares the number of digits in the serial number and fiscal year figures to determine whether the first and second C (information extraction targets) are both larger.

[0096] It is possible to comprehensively cover a variety of unstructured documents and appropriately identify information expression templates by specifying the information expression templates using the nine information expression grammars (row numbers #1 to #9) shown in the information expression template table 113. Then, by executing an information expression pattern generated using an information expression template specified by the information expression grammar and the support information type from this information expression template table 113 and performing pattern matching on unstructured documents, it is possible to accurately extract the information that the user is looking for from a variety of unstructured documents.

[0097] 7 is a flowchart illustrating a process in which the information expression structure analysis unit 120 analyzes information expressions contained in an unstructured document, and the information expression pattern generation unit 130 generates an information expression pattern using the results (hereinafter referred to as "information expression pattern generation process S700"). The information expression pattern generation process S700 will be described below with reference to this figure.

[0098] First, the information expression structure analysis unit 120 acquires an information expression from an unstructured document and registers it in the information expression group 111 (S701). For example, if the unstructured document is the unstructured document 300 illustrated in FIG. 3, the information expression structure analysis unit 120 acquires, for example, the area of ​​the company registration date 321 (the area including "Registration" and the date "07 / 16 / 2007") as an information expression. For example, the information expression acquired may be one that falls within a range specified by the user via the user device 3. stomach.

[0099] Next, the text information extraction unit 121 extracts text related to the extraction target from the information expression acquired in S701 (S702). For example, if the information expression is the company registration date 321 acquired from the unstructured document 300 illustrated in Fig. 3 and the extraction target is "07 / 16 / 2007", the text information extraction unit 121 extracts "Registration" and "07 / 16 / 2007" as text.

[0100] Next, the structural information extraction unit 122 extracts structural information related to the text extracted in S702 from the information expression acquired in S701 (S703). For example, if the information expression is the company registration date 321 acquired from the unstructured document 300 illustrated in FIG. 3, the structural information extraction unit 122 extracts the coordinates of the area surrounding each of the text "Registration" and the text "07 / 16 / 2007" extracted in S702 as structural information. The coordinates are expressed, for example, as a set of coordinates (xs, ys) of the upper left and coordinates (xe, ye) of the area surrounding the text.

[0101] Next, the information expression grammar identification unit 123 identifies an information expression grammar 1133 corresponding to the extracted text and structural information from the information expression template table 113 (S704). Details of this process (hereinafter referred to as "information expression grammar identification process S704") will be described later.

[0102] Next, the support information type identification unit 124 identifies the support information type corresponding to the extracted text and structure information from the support information type 1134 of the information expression template table 113 (S705). Details of this process (hereinafter referred to as "support information type identification process S705") will be described later.

[0103] Next, the information expression template search unit 131 of the information expression pattern generation unit 130 obtains an information expression template corresponding to the combination of the information expression grammar identified in S703 and the support information type identified in S704 from the information expression template table 113 (S706).

[0104] Next, the information expression element replacement unit 132 of the information expression pattern generation unit 130 replaces the bracketed notations in the information expression template acquired in S706 with the extraction target and the basis information to generate an information expression pattern (S707).

[0105] Fig. 8 is a flowchart for explaining details of the information expression grammar identification process S704 shown in Fig. 7. The information expression grammar identification process S704 will be explained below with reference to Fig. 8. Note that the following explanation will be given taking as an example a case where the information expression acquired in S701 of Fig. 7 (hereinafter referred to as "the information expression") is identified to which information expression grammar 1133 in the information expression template table 113 shown in Fig. 6 it corresponds.

[0106] The information expression grammar specification unit 123 first determines information to be extracted from the information expression (one or more C (information extraction target)) and information to be used for determining whether or not it is an extraction target (hereinafter referred to as “basis information”). (S801). It is not necessary to acquire the basis information. This information is received from the user via the user device 3, for example. For example, C (information extraction target) is displayed on the display device, and the user selects the target word by operating the mouse. The basis information is also displayed on the display device, and the user can click on the target word with the mouse or specify the target area. If the basis information is not visible, for example, the user can click only on C (information extraction target) with the mouse. Click.

[0107] Next, the information expression grammar identification unit 123 checks whether the number of C (information extraction targets) acquired in S801 is equal to or greater than the number of Cs. If there is only one C (information extraction target), then the answer is "Yes" (S802). ES) Proceed to S803, and if there are multiple C (information extraction targets), proceed to S820.

[0108] In S803, the information expression grammar identification unit 123 determines whether or not there is basis information. If there is no basis information obtained in S801 (S803: NO), the process proceeds to S804. If there is basis information obtained in S801 (S803: YES), the process proceeds to S805.

[0109] In S804, the information expression grammar identification unit 123 determines that the basis information is invisible information on the document (regular expressions, dictionary matches, meta information, control characters such as HTML tags, etc.), and determines that the information expression corresponds to the information expression grammar of row number "#1" in the information expression template table 113. Then, the information expression grammar identification process S704 ends.

[0110] In S805, the information expression grammar identification unit 123 determines whether the basis information is a range specification (S805). If the basis information is a range specification (S805: YES), the process proceeds to S806, where the information expression grammar identification unit 123 determines that the information expression corresponds to the information expression grammar of row number "#2" in the information expression template table 113. Thereafter, the information expression grammar identification process S704 ends. If the basis information is not a range specification (S805: NO), the process proceeds to S807.

[0111] In S807, the information expression grammar specification unit 123 determines whether the basis information is information (word) or information group (word group) that cannot be compared in magnitude with C (information extraction target) in information expression. It is determined whether or not the information expression grammar identification unit 123 is able to compare the magnitude of the basis information (word) or information group (word group). If the basis information is information (word) or information group (word group) that cannot be compared in magnitude (S807: NO), the process proceeds to S808, where the information expression grammar identification unit 123 determines that the information expression corresponds to the information expression grammar of row number "#3" in the information expression template table 113. Then, the information expression grammar identification process S704 ends. On the other hand, if the basis information is information (word) or information group (word group) that can be compared in magnitude (S807: YES), the process proceeds to S809, where the information expression grammar identification unit 123 determines that the information expression corresponds to the information expression grammar of row number "#4" in the information expression template table 113. Then, the information expression grammar identification process S704 ends.

[0112] In S820, the information expression grammar identification unit 123 determines whether or not there is basis information. If there is no basis information (S820: NO), the process proceeds to S821. If there is basis information (S820: YES), the process proceeds to S830.

[0113] In S821, the information expression grammar specification unit 123 determines whether C (information extraction target) is a numerical value or the like and whether it is a magnitude relation. It is judged whether the relationship between C and the value can be compared. If so (S821: YES), the process proceeds to S822, where the information expression grammar identification unit 123 determines that the information expression corresponds to the information expression grammar of row number "#8" in the information expression template table 113. Then, the information expression grammar identification process S704 ends. On the other hand, if the magnitude relationship cannot be compared (S821: NO), the process proceeds to S823, where the information expression grammar identification unit 123 determines that the information expression corresponds to the information expression grammar of row number "#5" in the information expression template table 113. Then, the information expression grammar identification process S704 ends.

[0114] In S830, the information expression grammar identification unit 123 determines whether the basis information is a range specification. If the basis information is a range specification (S830: YES), the process proceeds to S831, where the information expression grammar identification unit 123 determines that the information expression corresponds to the information expression grammar of row number "#6" in the information expression template table 113. Thereafter, the information expression grammar identification process S704 ends. On the other hand, if the basis information is not a range specification (S830: NO), the process proceeds to S833.

[0115] In S833, the information expression grammar specification unit 123 determines whether C (information extraction target) is a numerical value or the like and whether it is a magnitude relation. It is judged whether the relationship between C and the value can be compared. If so (S833: YES), the process proceeds to S834, where the information expression grammar identification unit 123 determines that the information expression corresponds to the information expression grammar of row number "#9" in the information expression template table 113. Then, the information expression grammar identification process S704 ends. On the other hand, if the magnitude relationship cannot be compared (S833: NO), the process proceeds to S835, where the information expression grammar identification unit 123 determines that the information expression corresponds to the information expression grammar of row number "#7" in the information expression template table 113. Then, the information expression grammar identification process S704 ends.

[0116] Fig. 9 is a flowchart for explaining the details of the support information type identification process S705 shown in Fig. 7. The support information type identification process S705 will be explained below with reference to Fig. 9. Note that the following explanation will be given taking as an example a case where the information expression acquired in S701 of Fig. 7 (hereinafter referred to as "the information expression") is identified to which support information type 1134 in the information expression template table 113 shown in Fig. 6 it corresponds.

[0117] First, the support information type specifying unit 124 acquires the information expression grammar specified in the information expression grammar specifying process S704 (S901).

[0118] Next, the support information type identification unit 124 determines whether the acquired information expression grammar is the information expression grammar of line number "#1" (S902). If the acquired information expression grammar is the information expression grammar of line number "#1" (S902: YES), proceed to S903, and if it is not the information expression grammar of line number "#1" (S902: NO), proceed to S910.

[0119] In S903, the support information type identification unit 124 determines whether the support information type of the information expression is a "regular expression." Specifically, the support information type identification unit 124 reads the regular expression dictionary of the various dictionaries 115 and determines whether C (information extraction target) matches the regular expression. The above determination is made by determining whether the support information type of the information expression is a "regular expression" (S903: YES), the support information type determination process S705 ends. On the other hand, if the support information type determination unit 124 determines that the support information type of the information expression is not a "regular expression" (S903: NO), the process proceeds to S904.

[0120] In S904, the support information type specification unit 124 determines whether the support information type of the information expression corresponds to "dictionary match." Specifically, the support information type specification unit 124 reads the word dictionary of the various dictionaries 115 and determines whether it matches C (information extraction target). As a result of the above determination, if the support information type of the information expression is determined to be a "dictionary match" (S904: YES), the support information type identification process S705 ends. On the other hand, if the support information type identification unit 124 determines that the support information type of the information expression is not a "dictionary match" (S904: NO), the process proceeds to S905.

[0121] In addition, there may be cases where the support information type for the information expression matches both "regular expression" and "dictionary match." In such cases, the user may be notified of this via the user device 3 and asked to select one of them, or the processing may be terminated without narrowing down the candidates to "regular expression" or "dictionary match."

[0122] In S905, the support information type identification unit 124 extracts meta information from the unstructured document and presents it to the user, prompting the user to select whether or not there is basis information. If the user determines that there is basis information in the presented meta information (S905: YES), the support information type identification unit 124 determines that the support information type of the information expression is "meta information" (page number, etc.) and selects the support information. On the other hand, if the user determines that the presented meta information does not contain any basis information (S905: NO), the process proceeds to S907.

[0123] In S907, the support information type identification unit 124 determines whether the unstructured document containing the information expression is written in HTML. If the unstructured document is written in HTML (S907: YES), the support information type identification unit 124 determines the support information type of the information expression to be "HTML structure," and the support information type identification process S705 ends. On the other hand, if the unstructured document is not written in HTML (S907: NO), the support information type identification unit 124 determines that there is no support information type that corresponds to the information expression (not applicable), and the support information type identification process S705 ends. Note that if it is determined that there is not applicable, the support information type identification unit 124 may prompt the user to input a regular expression or dictionary information, and determine the support information type of the information expression to be "regular expression" or "dictionary match."

[0124] In S910, the support information type identification unit 124 determines whether the information expression grammar acquired in S901 matches any of the information expression grammars with line numbers "#4," "#8," and "#9." If it matches any of the information expression grammars (S910: YES), the support information type identification unit 124 determines the support information type of the information expression to be "set," and the support information type identification process S705 ends. If it does not match any of the information expression grammars (S910: NO), proceed to S911.

[0125] In S911, the support information type specifying unit 124 determines whether C (information extraction target) and the basis information are It is determined whether the magnitude relationship can be compared for numerical values ​​or dates. If the magnitude relationship can be compared (S911: YES), the support information type identification unit 124 determines that the support information type of the information expression is "set," and the support information type identification process S705 ends. If the magnitude relationship cannot be compared (S911: NO), proceed to S912.

[0126] In S912, the support information type identification unit 124 determines whether the unstructured document containing the information expression is written in HTML. If the unstructured document is written in HTML (S912: YES), the support information type identification unit 124 determines the support information for the information expression to be "HTML structure," and the support information type identification process S705 ends. If the unstructured document is not written in HTML (S912: NO), the support information type identification unit 124 determines the support information to be "structure."

[0127] As described above, the document information management system 1 of the first embodiment can analyze the information expression structure contained in unstructured documents and generate information expression patterns to be used for extracting information from unstructured documents. The unstructured document management device 2 can then use the generated information expression patterns to efficiently extract information that the user wants to obtain from various unstructured documents with different formats.

[0128] [Second embodiment] 10 shows a schematic configuration of a document information management system 1 shown as the second embodiment. The information expression structure analysis device 100 in the document information management system 1 of the second embodiment further includes an information expression grammar identification support processing unit 140 in addition to the functions included in the information expression structure analysis device 100 of the first embodiment.

[0129] The information expression grammar identification support processing unit 140 receives information (C (information extraction target), grounds information, etc.) necessary for the information expression structure analysis unit 120 to identify the information expression grammar. Hereinafter, this information will be referred to as "identification support information." ) to assist in obtaining the certification.

[0130] Specifically, the information expression grammar identification support processing unit 140 receives the above information via the user device 3. A screen for receiving the information (hereinafter referred to as the "specific support information acquisition screen 1200") is presented to the user, and the above information is acquired from the user via the specific support information acquisition screen 1200. In this way, by presenting the specific support information acquisition screen 1200 and guiding the user to input the necessary information, it is possible to efficiently acquire specific support information even if the user does not have sufficient knowledge or experience in generating information expression patterns, for example.

[0131] 11 is a flowchart explaining the processing (hereinafter referred to as "specific support information acquisition processing S1100") that the information expression grammar specific support processing unit 140 performs when presenting the specific support information acquisition screen 1200 to the user and acquiring specific support information. The specific support information acquisition processing S1100 will be explained below with reference to this figure.

[0132] First, the information expression grammar specific support processing unit 140 presents the specific support information acquisition screen 1200 displaying an unstructured document to the user, and receives the designation of the first C (information extraction target) from the user. (S1101).

[0133] 12 shows an example of a specific support information acquisition screen 1200 presented to the user at this time. In the example shown in the figure, the user specifies the date "07 / 16 / 2007" 1211 as the first C (information extraction target). Therefore, the area is highlighted with a solid frame.

[0134] Returning to FIG. 11, the information expression grammar identification support processing unit 140 then receives from the user designation of the basis information used to extract the first C (information extraction target) or the second C (information extraction target) (S1102).

[0135] 13 shows an example of a specific support information acquisition screen 1200 presented to the user at this time. In this example, the user has specified the word "Registration" 1212 as the basis information, and therefore the area is highlighted with a dotted frame.

[0136] 11, the information expression grammar identification support processing unit 140 then determines whether or not the designation of the second C (information extraction target) has been accepted from the user in S1102 (S1103). If the designation of the second C (information extraction target) has been accepted (S1103: YES), the process proceeds to S1104, and if not accepted (S1103: NO), the process proceeds to S1121.

[0137] In S1104, the information expression grammar specific support processing unit 140 receives a selection of one of the support information types "HTML structure", "set", and "structure" from the user via the specific support information acquisition screen 1200.

[0138] 14 shows an example of a specific support information acquisition screen 1200 presented to the user at this time. In the example shown in the figure, a selection field 1213 for accepting the selection of one of "HTML structure," "set," and "structure" is displayed on the left side of the specific support information acquisition screen 1200. Note that, to help the user make a selection, for example, information expression patterns that are generated when each support information type is selected may be presented to the user.

[0139] Thereafter, the specific support information acquisition process S1100 ends, and the information expression grammar specific support processing unit 140 identifies an information expression grammar from the information expression template table 113 using the accepted support information type.

[0140] Returning to FIG. 11, in S1121, the information expression grammar specific support processing unit 140 accepts a selection of one of the support information types from the user via the specific support information acquisition screen 1200: “regular expression,” “dictionary match,” “meta information,” or “HTML structure.”

[0141] 15 shows an example of a specific support information acquisition screen 1200 presented to the user at this time. In the example shown in the figure, a selection field 1214 is displayed on the left side of the specific support information acquisition screen 1200, which accepts the selection of one of "regular expression," "dictionary match," "meta information," and "HTML structure." Note that, to help the user make a selection, for example, the information expression patterns that are generated when each support information type is selected may be presented to the user.

[0142] 11, the information expression grammar identification support processing unit 140 then determines whether or not the user has selected "regular expression" or "dictionary match" on the identification support information acquisition screen 1200 in Fig. 15 (S1122). If the user has selected "regular expression" or "dictionary match" (S1122: YES), the process proceeds to S1123; if not (S1122: NO), the process proceeds to S1124.

[0143] In S1123, the information expression grammar specific support processing unit 140 receives input of a regular expression or a dictionary from the user via the specific support information acquisition screen 1200.

[0144] 16 shows an example of a specific support information acquisition screen 1200 presented to the user at this time. In the example shown in the figure, an input field 1215 for accepting input of a regular expression or a dictionary is displayed on the left side of the specific support information acquisition screen 1200.

[0145] Thereafter, the specific support information acquisition process S1100 ends, and the information expression grammar specific support processing unit 140 identifies an information expression grammar from the information expression template table 113 using the received regular expression or dictionary contents.

[0146] Returning to Fig. 11, in S1124, the information expression grammar specific assistance processing unit 140 determines whether or not the user has selected "meta information" on the specific assistance information acquisition screen 1200 in Fig. 15. If the user has selected "meta information" (S1124: YES), the process proceeds to S1125, and if not (S1124: NO), the process proceeds to S1126.

[0147] In S1125, the information expression grammar specific assistance processing unit 140 receives the specification of meta information from the user via the specific assistance information acquisition screen 1200.

[0148] 17 shows an example of a specific support information acquisition screen 1200 presented to the user at this time. In the example shown in the figure, a selection field 1216 for accepting the selection of meta information is displayed on the left side of the specific support information acquisition screen 1200.

[0149] Thereafter, the specific support information acquisition process S1100 ends, and the information expression grammar specific support processing unit 140 identifies an information expression grammar from the information expression template table 113 using the received meta information.

[0150] Returning to FIG. 11, in S1126, the information expression grammar specific assistance processing unit 140 receives the specification of an HTML tag from the user via the specific assistance information acquisition screen 1200.

[0151] 18 shows an example of a specific support information acquisition screen 1200 presented to the user at this time. In the example shown in the figure, a selection field 1217 for accepting the selection of an HTML tag is displayed on the left side of the specific support information acquisition screen 1200.

[0152] Thereafter, the specific support information acquisition process S1100 ends, and the information expression grammar specific support processing unit 140 identifies the information expression grammar from the information expression template table 113 using the received HTML tag.

[0153] As described above, according to the document information management system 1 of the second embodiment, specific support information can be efficiently acquired from the user, and by using the specific support information to identify the information expression grammar and the support information type, an appropriate information expression template can be acquired and an information expression pattern can be efficiently generated.

[0154] [Third embodiment] 19 shows a schematic configuration of the document information management system 1 of the third embodiment. The information expression structure analysis device 100 of the document information management system 1 of the third embodiment further includes an information expression pattern verification unit 150 in addition to the functions of the information expression structure analysis device 100 of the first embodiment.

[0155] The information expression pattern verification unit 150 applies the information expression pattern generated by the information expression pattern generation unit 130 to an unstructured document, and presents the result to the user via the user device 3.

[0156] By using this function, the user can verify whether or not the desired information can be correctly extracted from an unstructured document using the information expression pattern generated by the information expression pattern generation unit 130. Also, for example, if there is a plurality of pieces of desired information, the user can verify whether or not each piece of information can be correctly extracted. If it turns out that the desired information cannot be extracted, the user can, for example, verify whether or not C (information extraction target) or root The basis information is reset and the information expression pattern is regenerated.

[0157] 20 is a flowchart illustrating the process (hereinafter referred to as "information expression pattern verification process S2000") that the information expression pattern verification unit 150 performs when verifying whether or not extraction information can be correctly extracted from an unstructured document using the information expression pattern generated by the information expression pattern generation unit 130. The information expression pattern verification process S2000 will be described below with reference to this figure.

[0158] First, the information expression pattern verification unit 150 acquires the information expression pattern generated by the information expression pattern generation unit 130 from the information expression pattern group 114 (S2001).

[0159] Next, the information expression pattern verification unit 150 extracts information from a predetermined unstructured document. All possible texts are extracted (S2002).

[0160] Next, the information expression pattern verification unit 150 inputs the text extracted in S2002 into the information expression pattern acquired in S2001, and checks whether the execution result of the information expression pattern is "TRUE" (S2003).

[0161] Next, the information expression pattern verification unit 150 generates a screen (hereinafter referred to as the "information expression pattern verification result display screen 2100") that highlights the text that has been determined to be "TRUE" along with the above-mentioned unstructured document, and presents the information expression pattern verification result display screen 2100 to the user via the user device 3.

[0162] 21 shows an information expression pattern verification result display screen 2100. As shown in the figure, the information expression pattern verification result display screen 2100 shown as an example displays the above-mentioned unstructured document, and text 2111 that is determined to be "TRUE" is highlighted by a dashed frame. By referring to the information expression pattern verification result display screen 2100, the user can efficiently verify whether or not the information expression pattern functions correctly.

[0163] Although one embodiment of the present invention has been described above, the present invention is not limited to the above embodiment. It goes without saying that various modifications can be made without departing from the spirit of the present invention. For example, the above-described embodiment has been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the described configurations. Furthermore, it is possible to add, delete, or replace part of the configuration of the above-described embodiment with other configurations.

[0164] Furthermore, the above-mentioned configurations, functional units, processing units, processing means, etc. may be partly or entirely implemented in hardware, for example, by designing them as integrated circuits. Furthermore, the above-mentioned configurations, functions, etc. may be implemented in software by a processor interpreting and executing a program that implements each function. Information such as programs, tables, and files that implement each function can be stored in a storage device such as a memory, a hard disk, or an SSD (Solid State Drive), It can be placed on recording media such as IC cards, SD cards, and DVDs.

[0165] Furthermore, the layout of the various functional units, processing units, and databases of each information processing device described above is merely an example, and the layout of the various functional units, processing units, and databases can be changed to an optimal layout in terms of the performance, processing efficiency, communication efficiency, etc. of the hardware and software that these devices are equipped with.

[0166] Furthermore, the configuration (schema, etc.) of the database that stores the various types of data described above can be flexibly changed from the perspective of efficient use of resources, improved processing efficiency, improved access efficiency, improved search efficiency, and the like. [Explanation of symbols]

[0167] 1 Document information management system, 2 Unstructured document management device, 21 Unstructured document management unit, 22 Information extraction unit, 23 Extracted information management unit, 24 Extracted information provision unit, 3 User device, 31 Various setting units, 32 Extracted information utilization unit, 100 Information expression structure analysis device, 110 Memory unit, 101 Extraction target information, 102 Basis information group, 111 Information expression group, 112 Information expression template group, 113 Information expression template table, 114 Information expression pattern group, 115 Various dictionaries, 120 Information expression structure analysis unit, 121 Text information extraction unit, 122 Structural information extraction unit, 123 Information expression grammar identification unit, 124 Support information type identification unit, 130 Information expression pattern generation unit, 131 Information expression template search unit, 132 Information expression component replacement unit, 140 Information expression grammar identification support processing unit, 150 Information expression pattern verification unit, S700 information expression pattern generation processing, S704 information expression grammar specification processing, S705 support information type specification processing, S1100 specific support information acquisition processing, 1200 specific support information acquisition screen, S2000 information expression pattern verification processing, 2100 information expression pattern verification result display screen< / header> < / header>

Claims

1. The information processing device includes a processor and a storage device, The storage device includes: Information representation, which is a manner of representing information in an unstructured document including text data and structural information; An extraction target is information that a user wants to extract from the unstructured document; grounds information used to determine whether or not to extract the extraction target from the information expression; Remember, The processor: extracting text data related to the extraction target from the information expression and extracting the structural information related to the extracted text data; The processor identifies which of the following (1) to (9) types of information expression grammar describes the information expression, based on the number of extracted text data, the presence or absence of the basis information, whether the basis information is within a specified range, and whether the basis information is information or a group of information that cannot be compared in magnitude with the information extraction target in the information expression: (1) The number of extracted text data is one, and the basis information is not acquired. (2) The number of extracted text data is one, the basis information is acquired, and the basis information is a range specification. (3) The number of extracted text data is one, the basis information is obtained, the basis information is not a range specification, and the basis information is information or a group of information that cannot be compared in size with the extracted target in terms of information expression. (4) The number of extracted text data is one, the basis information is obtained, the basis information is not a range specification, and the basis information is not information or a group of information that cannot be compared in size with the information extraction target in information expression. (5) The number of extracted text data is two or more, the basis information is not obtained, and the multiple extracted objects are information or groups of information that can be compared in terms of information expression. (6) The number of extracted text data is two or more, the basis information is not obtained, and the multiple extracted objects are not information or information groups that can be compared in terms of information expression. (7) The number of extracted text data is two or more, the basis information is acquired, and the basis information is a range specification. (8) The number of extracted text data is two or more, the basis information is obtained, the basis information is not a range specification, and the multiple extracted targets are information or groups of information that can be compared in terms of information expression. (9) The number of extracted text data is two or more, the basis information is obtained, the basis information is not a range specification, and the multiple extracted targets are not information or information groups that can be compared in terms of information expression. The processor identifies, based on the identified type of information expression grammar and type of structural information, which of the following (a) to (f) is a support information type, which is a classification based on the structure of the information expression: (a) The information expression grammar is the information expression grammar of (1) above, and is consistent with a regular expression. (b) The information expression grammar is the information expression grammar of (1) above, and the loaded dictionary words and the extraction target match. (c) The information expression grammar is the information expression grammar of (1), the loaded dictionary words do not match the extraction target, meta-information is extracted from the non-standard document, and the basis information exists in the extracted meta-information. (d) The information expression grammar is the information expression grammar of (1), the loaded dictionary word does not match the extraction target, meta-information is extracted from the non-standard document, the basis information does not exist in the extracted meta-information, and the non-standard document has been formatted with predetermined control characters embedded, or the information expression grammar is not any of (1), (4), (8), or (9), the extraction target and the basis information cannot be compared in terms of magnitude, and the non-standard document has been formatted with predetermined control characters embedded. (e) When the information expression grammar is any of (4), (8), or (9), or when the information expression grammar is none of (1), (4), (8), or (9), the extraction target and the basis information can be compared in terms of magnitude. (f) The information expression grammar is not any of (1), (4), (8), or (9), the extraction target and the basis information are not comparable in magnitude, and the non-standard document is not a document that has been formatted with specified control characters embedded. the storage device stores, for each of the information expression grammars, an information expression template table including information expression templates that are templates of information expression patterns that are program codes that realize a function of acquiring an extraction target from the unstructured document, the information expression templates being described in association with the support information types; the processor acquires, from the information expression template, the information expression template corresponding to the combination of the identified information expression grammar and the identified support information type; the processor generates the information expression pattern by replacing a symbol string of the information expression template with the extraction target and the basis information. Information representation structure analysis device.

2. 2. The information expression structure analysis device according to claim 1, The information expression grammar is a grammar expressing that the extraction target has the same meaning as a predetermined piece of information or a predetermined group of information; a grammar expressing that the position where the extraction target is described is within a predetermined area or a predetermined group of areas in the unstructured document; a grammar that expresses that the position where the extraction target is described in the unstructured document has a predetermined positional relationship with predetermined information or a predetermined group of information; a grammar that expresses that the position where the extraction target is described in the unstructured document has a predetermined relationship with predetermined information or a predetermined group of information; a grammar expressing that a position where a first extraction target is described has a predetermined positional relationship with a position where a second extraction target is described in the unstructured document; a grammar expressing that each of the plurality of extraction targets belongs to a predetermined area or a predetermined group of areas in an unstructured document; a grammar expressing that, in the unstructured document, each of the positions of the plurality of extraction targets has a predetermined positional relationship with a position where predetermined information or a group of information is described in the unstructured document; a grammar expressing that a first extraction target has a predetermined relationship with a second extraction target; and, a grammar expressing that the first extraction target and the second extraction target have a predetermined relationship with predetermined information or a group of information; is one of Information representation structure analysis device.

3. 2. The information expression structure analysis device according to claim 1, The support information type is at least one of a regular expression, a word dictionary, meta information, an HTML structure, a set of words, and an information expression structure. Information representation structure analysis device.

4. 2. The information expression structure analysis device according to claim 1, communicatively coupled to a user device; Accepting designation of the extraction target or the basis information via the user device while displaying a screen on which the non-standard document is described; receiving, via the user device, input of information corresponding to the support information type selected based on a result of receiving the extraction target or the basis information; Identifying the information expression grammar describing the information expression based on the received information. Information representation structure analysis device.

5. 2. The information expression structure analysis device according to claim 1, communicatively coupled to a user device; extracting text data that may be the extraction target from the unstructured document; inputting the acquired text data into the information expression pattern to check whether the execution result of the information expression pattern is true; generating a screen in which the text data for which the execution result is true is highlighted together with the unstructured document, and presenting the generated screen to the user via the user device; Information representation structure analysis device.

6. An information expression structure analysis method performed using an information processing device having a processor and a storage device, The storage device Information representation, which is a manner of representing information in an unstructured document including text data and structural information; An extraction target is information that a user wants to extract from the unstructured document; grounds information used to determine whether or not to extract the extraction target from the information expression; storing the the processor: extracting text data related to the extraction target from the information expression and extracting the structural information related to the extracted text data; a step of identifying which of the following (1) to (9) types of information expression grammar describing the information expression is, based on the number of extracted text data, the presence or absence of the basis information, whether the basis information is within a specified range, and whether the basis information is information or a group of information that cannot be compared in magnitude with the information extraction target in the information expression; (1) The number of extracted text data is one, and the basis information is not acquired. (2) The number of extracted text data is one, the basis information is acquired, and the basis information is a range specification. (3) The number of extracted text data is one, the basis information is obtained, the basis information is not a range specification, and the basis information is information or a group of information that cannot be compared in size with the extracted target in terms of information expression. (4) The number of extracted text data is one, the basis information is obtained, the basis information is not a range specification, and the basis information is not information or a group of information that cannot be compared in size with the information extraction target in information expression. (5) The number of extracted text data is two or more, the basis information is not obtained, and the multiple extracted objects are information or groups of information that can be compared in terms of information expression. (6) The number of extracted text data is two or more, the basis information is not obtained, and the multiple extracted objects are not information or information groups that can be compared in terms of information expression. (7) The number of extracted text data is two or more, the basis information is acquired, and the basis information is a range specification. (8) The number of extracted text data is two or more, the basis information is obtained, the basis information is not a range specification, and the multiple extracted targets are information or groups of information that can be compared in terms of information expression. (9) The number of extracted text data is two or more, the basis information is obtained, the basis information is not a range specification, and the multiple extracted targets are not information or information groups that can be compared in terms of information expression. a step of identifying, based on the identified type of information expression grammar and type of structural information, which of the following (a) to (f) is the type of support information, which is a classification based on the structure of the information expression; (a) The information expression grammar is the information expression grammar of (1) above, and is consistent with a regular expression. (b) The information expression grammar is the information expression grammar of (1) above, and the loaded dictionary words and the extraction target match. (c) The information expression grammar is the information expression grammar of (1), the loaded dictionary words do not match the extraction target, meta-information is extracted from the non-standard document, and the basis information exists in the extracted meta-information. (d) The information expression grammar is the information expression grammar of (1), the loaded dictionary word does not match the extraction target, meta-information is extracted from the non-standard document, the basis information does not exist in the extracted meta-information, and the non-standard document has been formatted with predetermined control characters embedded, or the information expression grammar is not any of (1), (4), (8), or (9), the extraction target and the basis information cannot be compared in terms of magnitude, and the non-standard document has been formatted with predetermined control characters embedded. (e) When the information expression grammar is any of (4), (8), or (9), or when the information expression grammar is none of (1), (4), (8), or (9), the extraction target and the basis information can be compared in terms of magnitude. (f) The information expression grammar is not any of (1), (4), (8), or (9), the extraction target and the basis information are not comparable in magnitude, and the non-standard document is not a document that has been formatted with specified control characters embedded. a step in which the storage device stores, for each of the information expression grammars, an information expression template table including information expression templates that are templates of information expression patterns that are program codes that realize a function of acquiring an extraction target from the unstructured document, the information expression templates being described in association with the support information type; the processor: acquiring, from the information expression template, the information expression template corresponding to the combination of the identified information expression grammar and the identified support information type; and generating the information expression pattern by replacing a symbol string of the information expression template with the extraction target and the basis information; This method performs information expression structure analysis.

7. 7. The information expression structure analysis method according to claim 6, The information expression grammar is a grammar expressing that the extraction target has the same meaning as a predetermined piece of information or a predetermined group of information; a grammar expressing that the position where the extraction target is described is within a predetermined area or a predetermined group of areas in the unstructured document; a grammar that expresses that the position where the extraction target is described in the unstructured document has a predetermined positional relationship with predetermined information or a predetermined group of information; a grammar that expresses that the position where the extraction target is described in the unstructured document has a predetermined relationship with predetermined information or a predetermined group of information; a grammar expressing that a position where a first extraction target is described has a predetermined positional relationship with a position where a second extraction target is described in the unstructured document; a grammar expressing that each of the plurality of extraction targets belongs to a predetermined area or a predetermined group of areas in an unstructured document; a grammar expressing that, in the unstructured document, each of the positions of the plurality of extraction targets has a predetermined positional relationship with a position where predetermined information or a group of information is described in the unstructured document; a grammar expressing that a first extraction target has a predetermined relationship with a second extraction target; and, a grammar expressing that the first extraction target and the second extraction target have a predetermined relationship with predetermined information or a group of information; is one of Information representation structure analysis method.

8. 7. The information expression structure analysis method according to claim 6, The support information type is at least one of a regular expression, a word dictionary, meta information, an HTML structure, a set of words, and an information expression structure. Information representation structure analysis device.

9. 7. The information expression structure analysis method according to claim 6, the information processing device is communicably connected to a user device; The information processing device, a step of accepting, via the user device, a designation of the extraction target or the basis information while displaying a screen on which the non-standard document is described; receiving, via the user device, input of information corresponding to the support information type selected based on a result of receiving the extraction target or the basis information; and identifying the information expression grammar describing the information expression based on the received information; The information expression structure analysis method further performs the above.

10. 7. The information expression structure analysis method according to claim 6, the information processing device is communicably connected to a user device; The information processing device, extracting text data that can be the extraction target from the unstructured document; inputting the acquired text data into the information expression pattern to check whether the execution result of the information expression pattern is true; and generating a screen in which the unstructured document and the text data for which the execution result is true are highlighted, and presenting the generated screen to the user via the user device; The information expression structure analysis method further performs the above.

Citation Information

Patent Citations

  • Grammar analysis of document visual structure

    JP2009500755A

  • System, method and program for creation of extraction rule

    JP2010262577A

  • System and method of automatic template generation

    US20200104354A1