Document processing device, program, and document processing method
The document processing device and method improve metadata extraction accuracy by identifying and categorizing character blocks within documents, addressing the challenge of incorrectly extracting similar document numbers in conventional technologies.
Patent Information
- Application Number
- JP2024570420
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-04-11
- Publication Date
- 2025-05-07
- Estimated Expiration
- 2043-04-11
AI Technical Summary
Conventional document processing technologies face challenges in accurately extracting metadata from documents, particularly when product model numbers are expressed in alphanumeric characters, leading to incorrect extraction of similar document numbers.
A document processing device and method that identify character blocks within a document, determine their types based on feature scores, and extract metadata by matching character strings with determined block types, using character block type determination information to calculate scores for each block type.
This approach enables more appropriate metadata extraction by accurately identifying and categorizing character blocks, reducing errors in extracting similar document numbers and improving overall metadata extraction accuracy.
Smart Images

Figure 0007672593000001 
Figure 0007672593000002 
Figure 0007672593000003
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a document processing device, a program, and a document processing method. [Background technology]
[0002] A technology has been proposed in the past that analyzes and scores each line of text in a document to see whether it satisfies the characteristics that are expected to be possessed by predefined metadata, and extracts the keyword with the highest score as metadata corresponding to the metadata identifier in the document (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 5380040 Summary of the Invention [Problem to be solved by the invention]
[0004] However, in the conventional technology, metadata is extracted only using the document data to be processed and a predefined dictionary, so for example, when a related product model number is characterized by being expressed in alphanumeric characters, a material number with a similar notation may be erroneously extracted.
[0005] Therefore, one or more aspects of the present disclosure aim to enable more appropriate extraction of metadata from document data. [Means for solving the problem]
[0006] A document processing device according to an aspect of the present disclosure includes: a character block identification unit that identifies a plurality of character blocks, each including a group of a plurality of characters, from a document represented by document data; a character block type determination unit that determines a character block type that is a type of each of the plurality of character blocks; and a metadata extraction unit that extracts, as metadata, from one character block, a character string that matches the extraction target character string in the character block type determined for one character block included in the plurality of character blocks, in accordance with metadata extraction condition information that indicates an extraction target character string that is a character string to be extracted from the one character block type. the character block type determination unit refers to character block type determination information indicating an added value for a feature possessed by each of the two or more character block types, calculates a score by adding up the added values when the plurality of characters have the feature, and determines, among the plurality of character blocks, a character block having the highest score in one character block type included in the two or more character block types as the one character block type; and determines, among the plurality of character blocks, a character block not determined to be any of the two or more character block types as a character block type other than the two or more character block types. It is characterized by:
[0007] A program according to an aspect of the present disclosure causes a computer to function as: a character block identification unit that identifies, from a document indicated by document data, a plurality of character blocks, each including a group of a plurality of characters; a character block type determination unit that determines a character block type that is a type of each of the plurality of character blocks; and a metadata extraction unit that extracts, as metadata, from one character block, a character string that matches the extraction target character string in the character block type determined for one character block included in the plurality of character blocks, in accordance with metadata extraction condition information that indicates an extraction target character string that is a character string to be extracted from the one character block. the character block type determination unit refers to character block type determination information indicating an added value for a feature possessed by each of the two or more character block types, calculates a score by adding up the added values when the plurality of characters have the feature, and determines, among the plurality of character blocks, a character block having the highest score in one character block type included in the two or more character block types as the one character block type; and determines, among the plurality of character blocks, a character block not determined to be any of the two or more character block types as a character block type other than the two or more character block types. It is characterized by:
[0008] A document processing method according to an embodiment of the present disclosure includes: The character block identification unit Identifying a plurality of character blocks, each of which includes a group of a plurality of characters, from the document represented by the document data; A character block type determination unit determining a character block type for each of the plurality of character blocks; The metadata extraction unit A document processing method for extracting, as metadata, from a character block, a character string that matches an extraction target character string in a character block type determined for a character block included in the plurality of character blocks, in accordance with metadata extraction condition information indicating an extraction target character string that is a character string to be extracted from each of the plurality of character block types, the one character block, the method comprising: The character block type determination unitThe method is characterized in that, by referring to character block type determination information indicating an additive value for a feature possessed by each of two or more character block types, a score obtained by adding up the additive values when the multiple characters have the feature is calculated for each of the two or more character block types, and a character block having the highest score in one character block type included in the two or more character block types among the multiple character blocks is determined to be the one character block type, and a character block not determined to be any of the two or more character block types among the multiple character blocks is determined to be a character block type other than the two or more character block types. Effect of the Invention
[0009] According to one or more aspects of the present disclosure, it becomes possible to extract metadata more appropriately. [Brief description of the drawings]
[0010] [Figure 1] 1 is a configuration diagram illustrating an outline of a configuration of a document processing device according to a first embodiment. [Diagram 2] 4 is a schematic diagram showing an example of character block type determination information according to the first embodiment. FIG. [Diagram 3] FIG. 11 is a schematic diagram showing an example of organization information stored in a character string DB. [Figure 4] 10 is a schematic diagram showing an example of product information stored in a character string DB. FIG. [Diagram 5] 10 is a schematic diagram showing an example of metadata extraction condition information. FIG. [Figure 6] FIG. 2 is a block diagram illustrating a schematic configuration of a PC. [Figure 7] 4 is a flowchart showing the operation of the document processing device according to the first embodiment. [Figure 8] 10 is a flowchart showing details of the operation of a character block type determining unit; [Figure 9] 1 is a schematic diagram showing an example of a document indicated in document data to be processed; [Figure 10] 13 is a schematic diagram showing an example of a character block identification result. FIG. [Figure 11] FIG. 11 is a configuration diagram illustrating an outline of a configuration of a document processing device according to a second embodiment. [Figure 12] FIG. 11 is a schematic diagram showing an example of character block type determination information in the second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] Embodiment 1 FIG. 1 is a diagram illustrating a configuration of a document processing device 100 according to the first embodiment. The document processing device 100 includes a character block type determination information storage unit 101, a character string DB (DataBase) 102, a metadata extraction condition information storage unit 103, a character block specification unit 104, a character block type determination unit 105, and a metadata extraction unit .
[0012] The character block type determination information storage unit 101 stores character block type determination information, which is information for determining the type of a character block. The character block type determination information indicates an added value for the characteristics possessed by each of two or more character block types. A character block is a group of characters identified from a document represented by document data, as described below. In other words, a character block includes a group of multiple characters.
[0013] FIG. 2 is a schematic diagram showing an example of character block type determination information according to the first embodiment. Character block type determination information 101a shown in FIG. 2 is information in a table format including an ID column 101b, a character block type column 101c, a feature column 101d, and a score addition value column 101e.
[0014] The ID column 101b stores ID (IDentification), which is identification information for identifying each row of the character block type determination information 101a.
[0015] The character block type column 101c stores the type of character block. The feature column 101d stores the features of the character block type. The score addition value column 101e stores an addition value that is a value to be added to the score of a character block type in the same line when the character blocks have the characteristics of the same line.
[0016] Character block type determination information 101a associates, for each character block type, a characteristic possessed by a character block with an additional value to be added as the score of the character block type when the character block has that characteristic.
[0017] 1, the character string DB 102 is a character string information storage unit that stores character string information indicating character strings that actually exist as metadata. For example, the character string DB 102 stores the names of departments that actually exist in organizations and product numbers that actually exist.
[0018] FIG. 3 is a schematic diagram showing an example of organization information stored in the character string DB 102. As shown in FIG. The organizational information 102a shown in FIG. 3 is information in a table format having an ID column 102b, a department column 102c, a section column 102d, and an abbreviation column 102e.
[0019] The ID column 102b stores an ID for identifying each row of the organization information 102a. The department column 102c stores the name of a department of an organization (for example, a company). The section column 102d stores the name of a section in the organization. A section is assumed to be a department to which an employee belongs. The abbreviation column 102e stores the abbreviation of the name of the section on the same row. From the organization information 102a, the name of the department, the name of the section, and the abbreviation of the section can be identified.
[0020] FIG. 4 is a schematic diagram showing an example of product information stored in the character string DB 102. As shown in FIG. The product information 102f shown in FIG. 4 is information in a table format having an ID column 102g and a product model number column 102h. The ID column 102g stores an ID for identifying each row of the product information 102f. The product model number column 102h stores the product model number. The product model number can be identified from the product information 102f.
[0021] 1, the metadata extraction condition information storage unit 103 stores metadata extraction condition information indicating conditions for extracting character strings to be extracted as metadata. The metadata extraction condition information indicates extraction target character strings, which are character strings to be extracted from each of a plurality of character block types.
[0022] FIG. 5 is a schematic diagram showing an example of metadata extraction condition information. The metadata extraction condition information 103a shown in FIG. 5 is information in a table format having an ID column 103b, a metadata type column 103c, a character block type column 103d to be extracted, and a character string column 103e to be extracted from the extraction target.
[0023] The ID column 103b stores an ID for identifying each row of the metadata extraction condition information 103a. The metadata type column 103c stores the metadata type, which is the type of metadata of the character string extracted from the extraction target. Extraction target character block type column 103d stores the type of character block of the extraction target character block from which a character string specified in character string to be extracted column 103e from the extraction target in the same row is extracted. The character string column 103e to be extracted from the extraction target stores extraction target character string information indicating the character string to be extracted from the character block of the character block type specified in the character block type column 103d to be extracted in the same row. The character string indicated by the extraction target character string information becomes the extraction target character string.
[0024] The metadata extraction condition information 103a specifies a character string to be extracted from a character block and the character block type of the character block from which the character string is to be extracted.
[0025] Returning to FIG. 1, character block identification unit 104 identifies character blocks from the document indicated by the document data. For example, if there is a frame, the character block identification unit 104 may identify the character string within the frame defined by the frame as a character block.
[0026] In addition, when there are no frame lines, the character block identification unit 104 may identify character blocks based on the width of the space between characters, the height of the space between lines, or the type of alignment between lines (e.g., centered, left aligned, right aligned). Specifically, the character block identification unit 104 determines that a character block is separated when the width of the space between characters is wider than a predetermined threshold value or when the width of the space between characters changes from the width of the space between characters in other locations. In addition, the character block identification unit 104 determines that a character block is cut off if the height of the space between lines is wider than a predetermined threshold value, or if the height of the space between lines changes from the height of the space between lines in other locations. Furthermore, character block identification unit 104 determines that a character block is broken when the alignment between lines changes. Then, the character block identification unit 104 identifies a character block, which is a group of characters, by appropriately combining the above judgments.
[0027] A character block type determining unit 105 determines the type of each of the plurality of character blocks. For example, the character block type determination unit 105 calculates a score for each character block type depending on whether the character block identified by the character block identification unit 104 satisfies the characteristics registered in the character block type determination information storage unit 101, and determines the character block type, such as title, author, main text, etc., based on the calculated score.
[0028] Specifically, character block type determination unit 105 refers to the character block type determination information and, when multiple characters contained in a character block have the characteristics indicated in the character block type determination information, calculates a score for each character block type by adding up the sums indicated in the character block type determination information. Then, character block type determination unit 105 determines, among the multiple character blocks, a character block with the highest score in one of the two or more character block types indicated in the character block type determination information as that one character block type. Furthermore, character block type determination unit 105 determines, among the multiple character blocks, a character block that is not determined to be any of the two or more character block types as a character block type other than the two or more character block types.
[0029] The metadata extraction unit 106 extracts, as metadata, a character string that matches the extraction target character string in the character block type from the character block determined to be of that character block type, in accordance with the metadata extraction condition information stored in the metadata extraction condition information storage unit 103. In addition, when the metadata extraction condition information indicates that a character string contained in the character string DB 102 is to be extracted as a character string to be extracted, the metadata extraction unit 106 extracts, as metadata, from the character block, a character string that matches the character string contained in the information stored in the character string DB 102.
[0030] Next, a hardware configuration of the document processing device 100 according to the first embodiment will be described. The document processing device 100 according to the first embodiment can be configured by a PC 10 as shown in FIG. The PC 10 includes a processor 11, a memory 12, an auxiliary storage device 13, a display 14, an input I / F (Interface) 15, and a communication I / F 16.
[0031] The processor 11 controls the entire document processing device 100. For example, the processor 11 is a central processing unit (CPU), a field programmable gate array (FPGA), or the like. The processor 11 may be a multiprocessor. The document processing device 100 may also include a processing circuit.
[0032] The memory 12 is a main storage device of the document processing device 100. For example, the memory 12 is a RAM (Random Access Memory) or the like. The auxiliary storage device 13 is a persistent storage device of the document processing device 100. For example, the auxiliary storage device 13 is a hard disk drive (HDD) or a solid state drive (SSD).
[0033] The display 14 displays various screen images. The input I / F 15 receives input of various instructions and is, for example, a keyboard and a mouse. The communication I / F 16 is an interface for transmitting and receiving data to and from other devices via a network such as the Internet, etc. For example, the communication I / F 16 is a network interface controller (NIC).
[0034] The above-described character block identification unit 104, character block type determination unit 105, and metadata extraction unit 106 can be realized by the memory 12 and the processor 11 that executes a program stored in the memory 12.
[0035] Such a program may be provided via a network, or may be provided in a form recorded on a recording medium, i.e., as a program product.
[0036] The character block type determination information storage unit 101 , the character string DB (DataBase) 102 , and the metadata extraction condition information storage unit 103 can be realized by the auxiliary storage device 13 .
[0037] Next, the operation of the document processing device 100 according to the first embodiment will be described. FIG. 7 is a flowchart showing the operation of the document processing device 100 according to the first embodiment. First, character block identification unit 104 identifies character blocks from the document indicated by the document data (S10). Then, for the identified character blocks, character block identification unit 104 generates character block information that associates the center coordinates of the character block, a character string, and its font information. The font information here only needs to include information used to determine the type of character block. For example, in the example shown in FIG. 2, the font information is "font size" and "bold".
[0038] Next, character block type determination unit 105 determines the character block type for each character block indicated by the character block information by referring to the information stored in character block type determination information storage unit 101 and character string DB 102 (S11). This process will be described in detail with reference to FIG. 8.
[0039] Next, the metadata extraction unit 106 refers to the information stored in the metadata extraction condition information storage unit 103 and the character string DB 102, and extracts metadata from the character string of the character block indicated by the character block information (S12). The metadata extraction unit 106 extracts, as metadata, a character string to be extracted from the extraction target indicated by the metadata extraction condition information from a character string included in a character block corresponding to the character block type to be extracted indicated by the metadata extraction condition information.
[0040] FIG. 8 is a flowchart showing the details of the operation of the character block type determining unit 105. First, character block type determination unit 105 refers to the character block type determination information stored in character block type determination information storage unit 101, and calculates a score for each character block indicated by the character block information (S20).
[0041] Next, for each character block type, character block type determination unit 105 determines the character block with the highest score calculated in step S20 as that character block type (S21). At this time, there may be character blocks to which multiple character block types are assigned. For character blocks to which no character block type is assigned, character block type determination unit 105 determines the character block type as "body."
[0042] An example of the operation of the document processing device 100 will be described with reference to FIGS. FIG. 9 is a schematic diagram showing an example of a document indicated by document data to be processed. FIG. 10 is a schematic diagram showing an example of a result of identifying character blocks for the document shown in FIG.
[0043] First, the character block identification unit 104 identifies character blocks B1 to B8 shown in FIG. 10 for the document shown in FIG. 9, and identifies the center coordinates, character string, and font information of the character block from each of the character blocks B1 to B8.
[0044] Next, character block type determining unit 105 determines the character block type for each of character blocks B1 to B8.
[0045] An example of a specific operation of the character block type determining unit 105 will be described below. Here, an example will be described in which the character block type is determined based on the character block type determination information 101a shown in FIG. 2, the organization information 102a shown in FIG. 3, and the product information 102f shown in FIG.
[0046] First, the score of the character block type "AUTHOR" of the character block B4 shown in FIG. 10 corresponds to the characteristics of ID1, ID2, and ID6 in FIG. 2, and is therefore 1.0+0.8+1.0+1 / 2=3.3. Furthermore, the score of the character block type "AUTHOR" of the character block B5 corresponds to the characteristics of ID1, ID2, ID5, and ID6 in FIG. 2, and is therefore 1.0+0.8+1.0+0.5+1 / 1=4.3.
[0047] Similarly, the character block type determining unit 105 calculates the scores of the character block type "AUTHOR" for the other character blocks B1 to B3 and B6 to B8. Then, the character block with the highest score for the character block type "AUTHOR" is character block B5, and therefore character block type determining unit 105 determines that character block B5 has a character block type of "AUTHOR."
[0048] Similarly, for the character block type "TITLE", character block type determining unit 105 calculates the scores for each of character blocks B1 to B8 and determines that character block B1 is of the character block type "TITLE".
[0049] Then, the character block type determining unit 105 determines that the character blocks B2 to B4 and B6 to B8 that are not determined to be either the character block type "AUTHOR" or the character block type "TITLE" are the character block type "BODY."
[0050] Next, an example of a specific operation of the metadata extraction unit 106 will be described. Here, as described above, it is assumed that the character block types of character blocks B1 to B8 shown in Figure 10 are determined based on character block type determination information 101a shown in Figure 2, organization information 102a shown in Figure 3, and product information 102f shown in Figure 4. Further, the following description will be given on the assumption that the metadata extraction condition information storage unit 103 stores metadata extraction condition information 103a.
[0051] In the metadata extraction condition information 103a, the character strings to be extracted from the extraction target, namely, "person's name," "organization name," and "date expression," are to be extracted from the character block type "AUTHOR." Therefore, the metadata extraction unit 106 extracts the person's name "Takahashi," the organization name "First Technology Division," and the date expression "2022 / 04 / 01" from the character block B5 by referring to the information stored in the character string DB102.
[0052] In addition, in the metadata extraction condition information 103a, the character string "product model number" to be extracted from the extraction target is to be extracted from the character block type "BODY", so the metadata extraction unit 106 extracts the product model number stored in the product information 102f from the character blocks B2 to B4 and B6 to B8. Here, "AA1234B" is extracted from the character block B7.
[0053] As described above, in the first embodiment, the character block type of the character block containing the character string to be extracted as metadata is identified by referring to the metadata extraction condition information 103a, so that the metadata can be extracted more appropriately.
[0054] Embodiment 2 In the above-mentioned first embodiment, the character string DB 102 that stores character strings that exist as metadata is used to determine the metadata, but the embodiment is not limited to this example. For example, document data from which metadata has already been extracted can be used to determine the metadata. This configuration will be described as the second embodiment.
[0055] FIG. 11 is a diagram illustrating a configuration of a document processing device 200 according to the second embodiment. The document processing device 200 includes a character block type determination information storage unit 201, a character string DB 102, a metadata extraction condition information storage unit 103, a character block specification unit 104, a character block type determination unit 205, a metadata extraction unit 206, and a document information DB 207.
[0056] The character string DB 102, metadata extraction condition information storage unit 103, and character block identification unit 104 of the document processing device 200 of embodiment 2 are similar to the character string DB 102, metadata extraction condition information storage unit 103, and character block identification unit 104 of the document processing device 100 of embodiment 1.
[0057] The document information DB 207 is a document information storage unit that stores document data from which metadata has already been extracted in association with the extracted metadata. The document data stored in the document information DB 207 is also referred to as accumulated document data.
[0058] Character block type determination information storage unit 201 stores character block type determination information. The character block type determination information in the second embodiment also includes information for determining the type of a character block based on a relationship with document data stored in document information DB 207.
[0059] FIG. 12 is a schematic diagram showing an example of character block type determination information according to the second embodiment. Character block type determination information 201a shown in FIG. 12 is information in a table format including an ID column 201b, a character block type column 201c, a feature column 201d, and a score addition value column 201e.
[0060] Here, ID column 201b, character block type column 201c, feature column 201d, and score addition value column 201e of character block type determination information 201a in embodiment 2 are the same as ID column 101b, character block type column 101c, feature column 101d, and score addition value column 101e of character block type determination information 101a in embodiment 1. However, in the character block type determination information 201a in embodiment 2, as shown by ID10 and ID11, information for determining the type of character block based on the relationship with the document data stored in the document information DB 207 is included.
[0061] Here, for example, the number of words appearing in both the document indicated by the document data to be processed and the document indicated by the document data stored in document information DB 207 is calculated, and if the number of words appearing in the document indicated by the document data to be processed that also appears in the document indicated by the document data stored in document information DB 207 out of the total number of words appearing in the document indicated by the document data to be processed exceeds a threshold value, it can be determined that these document data are similar. Note that the similarity of the document data may be determined by another method.
[0062] The character block type determination unit 205 calculates a score for each character block type depending on whether the character blocks identified by the character block identification unit 104 satisfy the characteristics registered in the character block type determination information storage unit 201, and determines the character block type, such as title, author, main text, etc., based on the calculated score. In the second embodiment, the character block type determining unit 205 also refers to the document data stored in the document information DB 207 to determine the type of the character block.
[0063] The metadata extraction unit 206 extracts, as metadata, from the corresponding character blocks, character strings that match the extraction target character strings in the character block type in accordance with the metadata extraction condition information stored in the metadata extraction condition information storage unit 103 .
[0064] Here, in the second embodiment, when the stored document data is similar to the document data from which metadata is to be extracted, the metadata extraction unit 206 extracts from the document data a character string that matches a specific character string contained in the stored document data.
[0065] Specifically, in the second embodiment, the metadata extraction unit 206 may refer to the document data and metadata stored in the document information DB 207, for example, search for document data similar to the document data to be processed in the document information DB 207, and extract metadata from the identified character block by determining whether or not the name of the creator of the similar document data is included in a character string.
[0066] As described above, according to the second embodiment, by using document data from which metadata has already been extracted, it is possible to extract metadata more appropriately. [Explanation of symbols]
[0067] 100, 200 document processing device; 101, 201 character block type determination information storage unit; 102 character string DB; 103 metadata extraction condition information storage unit; 104 character block identification unit; 105, 205 character block type determination unit; 106, 206 metadata extraction unit; 207 document information DB.
Claims
1. a character block identifying unit that identifies a plurality of character blocks, each of which includes a group of a plurality of characters, from a document represented by the document data; a character block type determination unit that determines a character block type that is a type of each of the plurality of character blocks; a metadata extraction unit that extracts, as metadata from one character block, a character string that matches the extraction target character string in a character block type determined for one character block included in the plurality of character blocks in accordance with metadata extraction condition information indicating an extraction target character string that is a character string to be extracted from each of the plurality of character block types, the character block type determination unit refers to character block type determination information indicating an added value for a feature possessed by each of the two or more character block types, calculates a score by adding up the added values when the plurality of characters have the feature, determines a character block having the highest score in one character block type included in the two or more character block types among the plurality of character blocks as the one character block type, and determines a character block not determined to be any of the two or more character block types among the plurality of character blocks as a character block type other than the two or more character block types. A document processing device comprising:
2. a document information storage unit that stores document data from which metadata has already been extracted as accumulated document data; The characteristics include a relationship with the stored document data.
2. The document processing device according to claim 1,
3. When a character string included in character string information is indicated as the extraction target character string in the metadata extraction condition information, the metadata extraction unit extracts a character string that matches the character string included in the character string information as metadata from the one character block.
3. The document processing device according to claim 1, wherein:
4. a document information storage unit that stores document data from which metadata has already been extracted as accumulated document data; When the stored document data is similar to the document data in which the character block is identified, the metadata extraction unit extracts a character string that matches a specific character string included in the stored document data from the document data in which the character block is identified.
2. The document processing device according to claim 1,
5. Computer, a character block identifying unit that identifies a plurality of character blocks, each of which includes a group of a plurality of characters, from a document represented by the document data; a character block type determination unit that determines a character block type that is a type of each of the plurality of character blocks; and a metadata extraction unit that extracts, according to metadata extraction condition information indicating an extraction target string, which is a string to be extracted from each of a plurality of character block types, a character string that matches the extraction target string in a character block type determined for one character block included in the plurality of character blocks, from the one character block as metadata; the character block type determination unit refers to character block type determination information indicating an added value for a feature possessed by each of the two or more character block types, calculates a score by adding up the added values when the plurality of characters have the feature, determines a character block having the highest score in one character block type included in the two or more character block types among the plurality of character blocks as the one character block type, and determines a character block not determined to be any of the two or more character block types among the plurality of character blocks as a character block type other than the two or more character block types. A program characterized by.
6. A character block identification unit identifies a plurality of character blocks, each of which includes a group of a plurality of characters, from a document represented by the document data; a character block type determination unit determines a character block type that is a type of each of the plurality of character blocks; a metadata extraction unit extracts, as metadata from one character block, a character string that matches an extraction target character string in a character block type determined for one character block included in the plurality of character blocks, in accordance with metadata extraction condition information indicating an extraction target character string that is a character string to be extracted from each of the plurality of character block types, the one character block comprising the plurality of character blocks, the metadata comprising: the character block type determination unit refers to character block type determination information indicating an added value for a feature possessed by each of the two or more character block types, calculates a score by adding up the added values when the plurality of characters have the feature, determines a character block having the highest score in one character block type included in the two or more character block types among the plurality of character blocks as the one character block type, and determines a character block not determined to be any of the two or more character block types among the plurality of character blocks as a character block type other than the two or more character block types. A document processing method comprising:
Citation Information
Patent Citations
Heat exchanger for hot air supply unit
JP1978080040A
Metadata extraction device and method therefor
JP2010102668A