Information Processing Apparatus, Information Processing Method, and Program

The information processing apparatus addresses the challenge of template selection in document creation by clustering documents and using decision trees to create templates, improving the efficiency of document generation.

JP7701114B2Active Publication Date: 2025-07-01NTT TECHNOCROSS CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022104030
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-06-28
Publication Date
2025-07-01
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

Existing document creation systems face challenges in determining the appropriate template to use from a plurality of options, making it difficult to efficiently create documents.

Method used

An information processing apparatus that clusters documents into multiple granularities using vectorization and decision trees, creating templates based on these clusters to assist in document creation.

Benefits of technology

Provides a technique for assisting in document creation by effectively selecting and utilizing appropriate templates, enhancing the efficiency of document generation processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007701114000001
    Figure 0007701114000001
  • Figure 0007701114000002
    Figure 0007701114000002
  • Figure 0007701114000003
    Figure 0007701114000003
Patent Text Reader

Abstract

To provide a technology for supporting in creating documents.SOLUTION: An information processing device according to one aspect includes: a cluster creation unit configured to cluster a plurality of document vectors into clusters with multiple granularities, the plurality of document vectors representing a plurality of respective documents or a plurality of respective masked documents in which a part of each of the plurality of respective documents is masked; a decision tree creation unit configured to create a decision tree for classifying the whole or a part of the plurality of documents or the plurality of masked documents in the descending order of the granularities, using a cluster ID of the clusters with the granularities as an objective variable and prescribed information related to the documents as an explanatory variable; and a template creation unit configured to create a template to be used for creating a target document based on the clustering results or the decision tree.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] Conventionally, techniques for creating document templates (skeletons) for the purpose of assisting in document creation, etc. have been known. For example, Non-Patent Document 1 discloses creating a template from a large number of legal documents by clustering for legal documents.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the prior art, when actually creating a document using a template, it has sometimes been difficult to determine which template to use from among a plurality of templates.

[0005] The present disclosure has been made in view of the above points, and provides a technique for assisting in document creation.

Means for Solving the Problems

[0006] An information processing apparatus according to one aspect of the present disclosure includes a cluster creation unit configured to cluster a plurality of document vectors each representing a plurality of documents or a plurality of masked documents in which each part of the plurality of documents is masked with a vector into clusters of a plurality of granularities, a decision tree creation unit configured to create a decision tree for classifying all or part of the plurality of documents or the plurality of masked documents, with the cluster ID of the cluster of the granularity as the target variable and predetermined information regarding the document as the explanatory variable, in descending order of the granularity, and a template creation unit configured to create a template used for creating a target document based on the clustering result or the decision tree.

Effect of the Invention

[0007] A technique for assisting in document creation can be provided.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Mode for Carrying Out the Invention

[0009] Hereinafter, an embodiment of the present invention will be described. In the following, a document creation apparatus 10 that can create a template (prototype) from an existing document set and create a document using the template will be described.

[0010] <Overall Configuration Example of Document Creation Apparatus 10> An overall configuration example of the document creation apparatus 10 according to this embodiment is shown in FIG. 1. As shown in FIG. 1, the document creation apparatus 10 according to this embodiment includes a template creation processing unit 101, a document creation processing unit 102, and an evaluation processing unit 103.

[0011] The document creation device 10 is realized by a general computer, a general-purpose server, or the like. The template creation processing unit 101, the document creation processing unit 102, and the evaluation processing unit 103 are realized, for example, by a process in which one or more programs installed in the document creation device 10 are executed by a processor (arithmetic unit) such as a CPU (Central Processing Unit). Further, the storage unit 104 is realized by a storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive), for example. Note that all or some of the functions of the template creation processing unit 101, the document creation processing unit 102, and the evaluation processing unit 103 may be provided by, for example, a cloud service or the like. Similarly, all or a part of the storage area of the storage unit 104 may be realized by, for example, cloud storage or the like.

[0012] The template creation processing unit 101 creates a template from a given document set. At this time, the template creation processing unit 101 clusters the given document set and creates a decision tree using the clustering result, and then creates a template from the clustering result or the decision tree.

[0013] The document creation processing unit 102 creates a document desired by the user using the template created by the template creation processing unit 101 (or a masked document in which a part of the documents included in the given document set is masked).

[0014] The evaluation processing unit 103 evaluates the validity of the clustering result and the like performed when creating the template.

[0015] The storage unit 104 stores various information (for example, a given document set, a masked document in which a part of the documents included in the document set is masked, a clustering result, a decision tree, conditions (branch conditions described later) held by nodes other than the leaves of the decision tree, a template, etc.).

[0016] Here, in the document creation apparatus 10 according to the present embodiment, there are a "template creation phase" in which the template creation processing unit 101 creates a template, a "document creation phase" in which the document creation processing unit 102 creates a document desired by the user, and an "evaluation phase" in which the evaluation processing unit 103 evaluates the validity of the clustering result and the like. Hereinafter, each of these phases will be described.

[0017] [Template Creation Phase] Hereinafter, the template creation phase will be described.

[0018] ·Example 1 First, Example 1 of the template creation phase will be described.

[0019] <Functional Configuration Example of Template Creation Processing Unit 101> A functional configuration example of the template creation processing unit 101 in Example 1 is shown in FIG. 2. As shown in FIG. 2, the template creation processing unit 101 in Example 1 includes a named entity extraction unit 201, a mask unit 202, a cluster creation unit 203, a decision tree creation unit 204, an end determination unit 205, and a template creation unit 206.

[0020] The named entity extraction unit 201 extracts named entities in each document included in the given document set. A named entity is an expression that represents a specific object with respect to a certain attribute (item), and typically includes proper nouns (personal names, organization names, place names, etc.), amounts, dates, times, and the like.

[0021] Here, an example of a document included in the given document set is shown in FIG. 3. In FIG. 3, as an example, an expense expenditure inquiry form, which is one of the approval documents, is shown. As shown in FIG. 3, the document contains various attributes (items). In the document shown in FIG. 3, "application date", "applicant", "department", "subject", "implementation purpose" are included as attributes, and "implementation content (purpose)", "order content and quantity", "contract period", "order recipient", "reason for selecting the contractor", "contract amount", "payment method", "authority and responsibility regulations", "others" in the text are included as attributes.

[0022] Note that in the example shown in FIG. 3, each attribute is described in the entry field where the text is entered, but this is just an example. For example, there may be an entry field or the like for entering the character string of each attribute for each attribute. Also, in FIG. 3, an expense expenditure inquiry form, which is one of the approval documents, is shown, but this is just an example, and this embodiment is applicable to various types of documents.

[0023] The mask part 202 replaces the specific expression existing in the document where the specific expression is extracted by the specific expression extraction part 201 with a predetermined character string. This means masking the specific expression in the document with a certain predetermined character string. As a result, a masked document in which the specific expressions in each document included in the document set are masked with a predetermined character string is created, and as a result, a masked document set composed of these masked documents is obtained.

[0024] Here, substitution examples of specific expressions are shown in FIG. 4. As shown in FIG. 4, according to the attributes of the specific expressions, the specific expressions are replaced with character strings corresponding to the attributes. For example, when the attribute of a certain specific expression is "personal name", the specific expression is replaced with "PERSON". Similarly, for example, when the attribute of a certain specific expression is "company name", the specific expression is replaced with "COMPANY". Similarly, for example, when the attribute of a certain specific expression is "organization name", the specific expression is replaced with "ORGANIZATION". In this way, the specific expressions extracted by the specific expression extraction unit 201 are replaced with character strings according to their attributes. Note that the attributes of the specific expressions shown in FIG. 4 and the corresponding replacement character strings are just examples and are not limited thereto. For example, they may be replaced with symbols or the like.

[0025] Note that the extraction of specific expressions and the replacement of specific expressions can be realized by known methods. For example, for the extraction of specific expressions, the extended specific expressions described in Reference 1 can be used.

[0026] The cluster creation unit 203 clusters each masked document included in the masked document set by hierarchical clustering. At this time, the cluster creation unit 203 classifies each masked document into clusters of multiple granularities based on the similarity of the masked documents. Note that the clustering result of hierarchical clustering is also called a dendrogram. Here, the granularity of a cluster refers to the number of masked documents classified into that cluster. The coarser the granularity, the more masked documents are included in one cluster. Conversely, the finer the granularity, the fewer masked documents are included in one cluster. Hereinafter, as an example, it is assumed that the clustering is performed into two types: a "coarse cluster" and a "fine cluster" which has a finer granularity than the "coarse cluster". However, this is just an example, and the granularity of the cluster may be three or more.

[0027] Specifically, the cluster creation unit 203 clusters each masked document into a "coarse cluster" and a "fine cluster" according to the following procedures 11 to 13.

[0028] Step 11: First, the cluster creation unit 203 vectorizes each masked document. Hereinafter, the vectorized masked document will be referred to as a "masked document vector" or simply as a "document vector" when there is no risk of misunderstanding. Note that the vectorization of the masked document may be performed by a known method such as tf-idf, Doc2Vec, or Word2Vec. Further, instead of vectorizing the entire masked document to create a document vector, for example, only the title or subject of the masked document may be vectorized to create a document vector, or only the main text may be vectorized to create a document vector, or a summary of the masked document may be created by a known summary creation method and then the summary may be vectorized to create a document vector. Or, for example, when the original document is an approval document or the like, metadata such as the approval route, approval authority, or some ID code associated with the original document of the masked document may be vectorized to create a document vector.

[0029] Step 12: Next, the cluster creation unit 203 clusters each document vector by hierarchical clustering using a threshold value for clustering into rough clusters (hereinafter referred to as the first threshold value). As a result, each document vector (i.e., each masked document) is classified into a rough cluster, and as a result, the cluster ID of the rough cluster is assigned to each document vector. Hereinafter, assuming that the number of rough clusters is N, the cluster IDs of the rough clusters are "A1", "A2", ···, "AN".

[0030] Step 13: Then, the cluster creation unit 203 further clusters each document vector classified into a rough cluster by hierarchical clustering for each rough cluster using a threshold value for clustering into fine clusters (hereinafter referred to as the second threshold value). As a result, each document vector is further classified into a fine cluster, and as a result, the cluster ID of the fine cluster is assigned to each document vector. Hereinafter, assuming that the number of fine clusters is n1 + n2 + ··· + n NAs the cluster IDs of the fine-grained clusters, use "a11", "a12", ···, "a1n1", "a21", "a22", ···, "a2n2", ···, "aN1", "aN2", ···, "aNn N ".

[0031] Note that as the similarity between document vectors, for example, Euclidean distance, cosine similarity, etc. can be adopted. Also, the first threshold is a threshold for roughly classifying each document vector, and the second threshold is a threshold for finely classifying each document vector. For example, when Euclidean distance is adopted as the similarity between document vectors, the first threshold < the second threshold is satisfied. If the coarser cluster contains more document vectors than the finer cluster, the number of document vectors included in each cluster is not particularly limited, but it is preferable that the coarser cluster contains about 10 times more document vectors than the finer cluster. As a specific example, when the number of masked documents included in the masked document set is about 1000, it is preferable that the coarser cluster contains about 100 document vectors and the finer cluster contains about 10 document vectors.

[0032] An example of the hierarchical clustering result (dendrogram) based on the similarity between document vectors by the above-mentioned cluster creation unit 203 is shown in FIG. 5. In FIG. 5, a case is shown where each masked document included in the masked document set is classified into a coarser cluster (cluster IDs "A1" to "AN") and a finer cluster (cluster IDs "a11" to "a1n1", ···, cluster IDs "aN1" to "aNn N ").

[0033] The decision tree creation unit 204 creates decision trees (overall decision tree and partial decision tree) from the masked document vectors using the clustering result by the cluster creation unit 203, the internal information, and the external information of the document. At this time, the decision tree creation unit 204 uses the clustering result when clustering into rough clusters as teacher data, the cluster ID of the rough cluster as the target variable, and the internal information and the external information as explanatory variables, and creates a decision tree (hereinafter, this decision tree is referred to as the "overall decision tree") by a known decision tree creation algorithm (for example, CART, etc.). Further, for each rough cluster, the decision tree creation unit 204 uses the clustering result when further clustering the rough cluster into fine clusters as teacher data, the cluster ID of the fine cluster as the target variable, and the internal information and the external information as explanatory variables, and creates a decision tree (hereinafter, this decision tree is referred to as the "partial decision tree") by a known decision tree creation algorithm.

[0034] The internal information is the internal information of the document, for example, words or word sequences included in the document, words or word sequences included in a certain attribute of the document, keywords included in a certain attribute of the document, etc. For example, specific examples of words included in the document include "confidential" and "contract", and specific examples of word sequences included in the document include "confidentiality maintenance". Further, for example, a specific example of a word included in the document attribute "order destination" is "company", and specific examples of word sequences included in the document attribute "order destination" include "corporation". Similarly, for example, a specific example of a keyword included in the document attribute "implementation content (purpose)" is "consumable purchase". In addition to these, the internal information may include a document structure such as a pattern of an attribute including numbers, symbols, etc.

[0035] The external information is the external information of the document (or information attached to the document), for example, words or word sequences included in regulations related to the document. For example, when the document is an approval document, it is words or word sequences included in the clause in which the approval authority of the document is defined. Specific examples of such words include "intellectual property", and specific examples of word sequences include "intellectual property management".

[0036] Specifically, the decision tree creation unit 204 creates the overall decision tree and the partial decision tree according to the following procedures 21 to 22.

[0037] Procedure 21: First, the decision tree creation unit 204 creates an overall decision tree from the masked document vectors, using the clustering result when clustering into rough clusters as the training data, the cluster IDs "A1" to "AN" of the rough clusters as the target variable, and the internal information and external information as the explanatory variables. As a result, an overall decision tree for classifying each document vector into a rough cluster is created in the same manner as the cluster creation unit 203. Here, an example of the overall decision tree is shown in FIG. 6. As shown in FIG. 6, the overall decision tree is represented by a graph structure having the cluster IDs of the rough clusters as the leaves (leaf nodes) and the conditions (branch conditions) regarding the explanatory variables as the nodes (nodes other than the leaves). Note that specific examples of the branch conditions include, for example, "the document contains the word sequence 'confidentiality maintenance'", "the word 'company' is included in the attribute 'order destination'", "the word 'intellectual property' is included in the clause where the approval authority of the document is defined", and the like.

[0038] Procedure 22: Next, for each coarse cluster, the decision tree creation unit 204 creates a partial decision tree from the masked document vectors classified into the coarse cluster, using the clustering result when the coarse cluster is further clustered into fine clusters as the training data, the cluster ID of the fine cluster as the target variable, and the internal variables and external variables as the explanatory variables. At this time, the decision tree creation unit 204 may create a partial decision tree from the masked document vectors clustered into the coarse cluster by the cluster creation unit 203, or may create a partial decision tree from the masked document vectors classified into the leaf node corresponding to the coarse cluster by the overall decision tree created in Procedure 21. As a result, a partial decision tree for further classifying each document vector classified into a certain coarse cluster into finer clusters is created, similar to the cluster creation unit 203. Here, FIG. 7 shows an example of a partial decision tree for classifying the masked document vectors classified into the coarse cluster with the cluster ID "A1" into the fine clusters with the cluster IDs "a11" to "a1n1". As shown in FIG. 7, the partial decision tree is represented by a graph structure having the cluster IDs "a11" to "a1n1" as leaves (leaf nodes) and the branching conditions regarding the explanatory variables as nodes (nodes other than leaves).

[0039] In this way, since an overall decision tree for explaining the clustering result for the coarse clusters and a partial decision tree for explaining the clustering result when each coarse cluster is further clustered into fine clusters are obtained, for example, when creating a desired document, the user can select and use an appropriate template with reference to the overall decision tree and the partial decision tree.

[0040] Note that the algorithm for creating the overall decision tree and the algorithm for creating the partial decision trees may be the same algorithm or different algorithms. Also, when the number of cluster granularities is three or more, decision trees may be created in order with the cluster ID of the cluster of that granularity as the target variable in descending order of granularity. Specifically, for example, if there are a first granularity, a second granularity, and a third granularity as the cluster granularities, and they are arranged in descending order in this order, after creating a decision tree (overall decision tree) with the cluster ID of the cluster of the first granularity as the target variable, for each cluster of the first granularity, a decision tree (partial decision tree) is created with the cluster ID when that cluster is further clustered into clusters of the second granularity as the target variable, and then, for each cluster of the second granularity, a decision tree (partial decision tree) is created with the cluster ID when that cluster is further clustered into clusters of the third granularity as the target variable.

[0041] The end determination unit 205 determines whether to end the clustering, creation of the overall decision tree, and creation of the partial decision trees. At this time, when the end determination unit 205 satisfies a predetermined end condition, it determines to end the clustering, creation of the overall decision tree, and creation of the partial decision trees, and otherwise determines not to end. Note that when it is determined by the end determination unit 205 not to end, for example, after changing the first threshold value and the second threshold value, the process starts over from the clustering by the cluster creation unit 203. How to change the first threshold value and the second threshold value can be arbitrarily determined as appropriate, but for example, it is conceivable to add a predetermined value determined in advance to the first threshold value and the second threshold value, or subtract from the first threshold value and the second threshold value.

[0042] However, it is not limited to this. When it is determined by the end determination unit 205 not to end, for example, the process may start over from the creation of the overall decision tree and the partial decision trees. Or, for example, the similarity index may be changed (for example, changed from the Euclidean distance to the cosine similarity, etc.).

[0043] Here, as the termination condition, various conditions can be used which determine to end clustering and the creation of the overall decision tree and partial decision trees when the clustering result is the same as or similar to the classification results of the overall decision tree and partial decision trees, and determine not to end otherwise. For example, "the number of masked document vectors for which the clustering result is the same as the classification results of the overall decision tree and partial decision trees is equal to or greater than a predetermined threshold", "the ratio of masked document vectors for which the clustering result is the same as the classification results of the overall decision tree and partial decision trees is equal to or greater than a predetermined ratio", "a statistical value regarding the classification results of the overall decision tree and partial decision trees (for example, the average value of the ratio of masked document vectors for which the classification result of each leaf is the same as the clustering result, etc.) is equal to or greater than a predetermined threshold", and the like.

[0044] When the template creation unit 206 determines that the clustering, creation of the overall decision tree, and creation of the partial decision trees have been completed by the end determination unit 205, the template creation unit 206 creates a template from the clustering result by the cluster creation unit 203. For example, for each fine-grained cluster, the template creation unit 206 creates, as a template, the masked document corresponding to the masked document vector (hereinafter also referred to as the "representative document vector") that represents the fine-grained cluster. Here, as the representative document vector, the masked document vector closest to the centroid of the fine-grained cluster can be mentioned, but it is not limited to this. For example, a masked document vector randomly selected from the fine-grained cluster may be used as the representative document vector, or a masked document vector selected according to a predetermined distribution may be used as the representative document vector. Also, the masked document itself corresponding to the representative document vector may be used as the template, or something obtained by adding a certain attribute to the masked document, or deleting a certain attribute from the masked document may be used as the template, or something obtained by adding some sentences to the masked document, or deleting some sentences from the masked document may be used as the template. For example, it is conceivable to delete from the masked document used as the template the attributes other than those common to the masked documents corresponding to the document vectors belonging to the fine-grained cluster, or to delete the sentences other than the common sentences. In addition to this, depending on the type of document, predetermined attributes may be added or predetermined sentences may be added as appropriate.

[0045] As a result, a template is created for each fine-grained cluster, and a template set composed of these templates is obtained. As will be described later, by using a desired template in the template set and changing the masked part (that is, the masked part of the specific expression) in the template to a desired character string, a desired document can be created.

[0046] <Template Creation Process> Hereinafter, the template creation process in the first embodiment will be described with reference to FIG. 8.

[0047] Step S101: First, the named entity extraction unit 201 extracts named entities in each document included in the given document set.

[0048] Step S102: Next, the masking unit 202 replaces the named entity existing in the document where the named entity was extracted in step S101 with a predetermined character string for the named entity extracted in step S101 above. Thereby, a masked document set is obtained.

[0049] Note that in steps S101 to S102 above, a masked document set was created from the document set. However, for example, when a masked document set is given to the document creation device 10 instead of the document set, steps S101 to S102 above are unnecessary.

[0050] Step S103: Next, the cluster creation unit 203 clusters each masked document included in the masked document set into clusters of multiple granularities by hierarchical clustering. That is, the cluster creation unit 203, for example, clusters each masked document into a coarse cluster by the above-described procedure 11 to procedure 13, and then further clusters it into a fine cluster.

[0051] Step S104: Next, the decision tree creation unit 204 creates an overall decision tree from the masked document vectors by the above-described procedure 21 using the clustering result (clustering result when clustering into coarse clusters) of step S103 above, internal information, and external information.

[0052] Step S105: Next, for each coarse cluster, the decision tree creation unit 204 creates a partial decision tree from the masked document vectors classified into the coarse cluster by the above-described procedure 22 using the clustering result (clustering result when further clustering the coarse cluster into fine clusters) of step S103 above, internal information, and external information.

[0053] Step S106: Next, the end determination unit 205 determines whether a predetermined end condition is satisfied. If it is determined that the end condition is not satisfied, the template creation processing unit 101 changes the first threshold value and the second threshold value, and then returns to step S103 to start over from clustering. On the other hand, if it is determined that the end condition is satisfied, the template creation processing unit 101 proceeds to step S107.

[0054] Step S107: Then, the template creation unit 206 uses the clustering result of step S103 above (the clustering result when clustering into fine clusters) to create, for each fine cluster, a masked document corresponding to the representative document vector of the fine cluster as a template.

[0055] Note that the above template creation process is an example and can be changed as appropriate. For example, as described above, if it is determined that the predetermined end condition is not satisfied, the template creation processing unit 101 may start over from creating a decision tree (the overall decision tree and the partial decision tree).

[0056] Also, in the above template creation process, after creating the decision trees (the overall decision tree and partial decision trees), the end determination unit 205 performs an end determination. Instead of or in addition to this, for example, during the creation of the decision tree, a determination may be made as to whether to recreate the decision tree. In this case, for example, when it is determined to recreate the decision tree, after changing the branching conditions, etc. of the decision tree, the creation of the decision tree may be restarted from the beginning, or restarted from one level above the decision tree. Various conditions can be cited as conditions for determining whether to recreate the decision tree. For example, "the number of masked document vectors with classification results different from the clustering results at the current time is equal to or greater than a predetermined threshold", "the ratio of masked document vectors with classification results different from the clustering results at the current time is equal to or greater than a predetermined ratio", "the statistical value regarding the classification result of the decision tree at the current time (for example, the average value of the ratio of masked document vectors whose classification results of each leaf are the same as the clustering results, etc.) is less than a predetermined threshold", and the like.

[0057] ·Example 2 Next, Example 2 of the template creation phase will be described. In Example 2 of the template creation phase, the differences from Example 1 will be described, and the description of the components that may be the same as those in Example 1 will be omitted.

[0058] <Functional configuration example of the template creation processing unit 101> A functional configuration example of the template creation processing unit 101 in Example 2 is shown in FIG. 9. As shown in FIG. 9, the template creation processing unit 101 in Example 2 includes a unique expression extraction unit 201, a mask unit 202, a cluster creation unit 203, an overall decision tree creation unit 204A, a partial decision tree creation unit 204B, an end determination unit 205, and a template creation unit 206.

[0059] The overall decision tree creation unit 204A creates an overall decision tree from the masked document vectors using the clustering result (the clustering result when clustering into rough clusters) by the cluster creation unit 203, the internal information and external information of the document.

[0060] The partial decision tree creation unit 204B creates a partial decision tree from the masked document vectors classified into the coarse clusters by using the clustering result by the cluster creation unit 203 (the clustering result when the coarse clusters are further clustered into fine clusters) for each coarse cluster, and the internal information and external information of the documents.

[0061] As described above, the template creation processing unit 101 in the second embodiment includes an overall decision tree creation unit 204A that creates an overall decision tree and a partial decision tree creation unit 204B that creates each partial decision tree, instead of the decision tree creation unit 204. In this case, step S104 in FIG. 8 is executed by the overall decision tree creation unit 204A, and step S105 is executed by the partial decision tree creation unit 204B. Other points are the same as those in the first embodiment.

[0062] · Embodiment 3 Next, an example of the template creation phase in the third embodiment will be described. In the third embodiment of the template creation phase, differences from the first embodiment will be described, and descriptions of components that may be the same as those in the first and second embodiments will be omitted.

[0063] <Functional configuration example of the template creation processing unit 101> A functional configuration example of the template creation processing unit 101 in the third embodiment is shown in FIG. 10. As shown in FIG. 10, the template creation processing unit 101 in the third embodiment includes, in the same manner as in the first embodiment, a unique expression extraction unit 201, a mask unit 202, a cluster creation unit 203, a decision tree creation unit 204, an end determination unit 205, and a template creation unit 206, but the function of the template creation unit 206 is different from that in the first embodiment.

[0064] In Example 3, when the template creation unit 206 is determined by the end determination unit 205 to have completed clustering and the creation of the overall decision tree and partial decision trees, it creates a template from the classification results of each partial decision tree. For example, the template creation unit 206 creates, for each leaf (leaf node) of each partial decision, a masked document corresponding to the representative document vector of the masked document vectors classified into that leaf as a template. Here, examples of the representative document vector in Example 3 include the document vector of a masked document that has all the attributes included in each branch condition from the root of the decision tree to that leaf among the masked document vectors classified into that leaf. However, it is not limited to this. For example, a masked document vector randomly selected from that leaf may be used as the representative document vector, or a masked document vector selected according to a predetermined distribution may be used as the representative document vector.

[0065] Note that although templates are created for each leaf of each partial decision tree, it is not limited to this. Instead of or in addition to this, for example, templates may be created for each leaf of the overall decision tree.

[0066] In this way, the template creation unit 206 in Example 3 creates a template from the decision tree. In this case, in step S107 of FIG. 8, the template creation unit 206 creates a template from the decision tree (the overall decision tree, partial decision trees, or both). Other aspects are the same as in Example 1 or 2.

[0067] [Document Creation Phase] Hereinafter, the document creation phase will be described.

[0068] ·Example 1 First, Example 1 of the document creation phase will be described.

[0069] [Functional Configuration Example of Document Creation Processing Unit 102] The functional configuration example of the document creation processing unit 102 in Embodiment 1 is shown in FIG. 11. As shown in FIG. 11, the document creation processing unit 102 in Embodiment 1 includes an index creation unit 301, a search unit 302, and a document creation unit 303.

[0070] When an attribute to be indexed (hereinafter also referred to as "target attribute") is given, the index creation unit 301 creates an index of a template set related to the target attribute. The target attribute may be given by a user or the like, or a predetermined attribute may be given. Note that the index may be created by any known method. For example, an inverted index may be created using masked document vectors corresponding to the templates constituting the template set.

[0071] When a search condition related to the target attribute is given, the search unit 302 searches for a template that satisfies the search condition from the template set using the index created by the index creation unit 301. The search condition may be given by a user or the like, or may be given from another program, system, or the like.

[0072] The document creation unit 303 creates a desired document (hereinafter also referred to as "target document") using the template searched by the search unit 302 and a character string to be set at the masked portion of the template (hereinafter also referred to as "setting character string"). That is, the document creation unit 303 creates the target document by setting the setting character string corresponding to the masked portion for the masked portion of the template. Thereby, the target document desired by the user can be obtained. The setting character string may be given by a user or the like, or may be given from another program, system, or the like.

[0073] <Document Creation Process> Hereinafter, the document creation process in Example 1 will be described with reference to FIG. 12. Here, step S201 in FIG. 12 is performed in advance before creating the target document, and steps S202 to S203 are performed each time the target document is created. However, when recreating the index, step S201 is performed at an appropriate timing.

[0074] Step S201: When the target attribute is given, the index creation unit 301 creates an index of the template set related to the target attribute.

[0075] Step S202: When the search condition related to the target attribute is given, the search unit 302 uses the index created in step S201 above to search for a template that satisfies the search condition from the template set.

[0076] Step S203: Then, the document creation unit 303 creates the target document by setting the given setting string for the masked part of the template searched in step S202 above.

[0077] Note that the above document creation process is an example and can be changed as appropriate. For example, as described in Example 3 to be described later, when no index is created, step S201 is unnecessary.

[0078] ·Example 2 Next, Example 2 of the document creation phase will be described. Note that in Example 2 of the document creation phase, the differences from Example 1 will be described, and the description of the components that may be the same as those in Example 1 will be omitted.

[0079] <Functional configuration example of document creation processing unit 102> A functional configuration example of the document creation processing unit 102 in Example 2 is shown in FIG. 13. As shown in FIG. 13, the document creation processing unit 102 in Example 2 includes an index creation unit 301, a search unit 302, and a document creation unit 303, similar to Example 1, but the functions of these units are different from those in Example 1.

[0080] In Example 2, when a target attribute is given, the index creation unit 301 creates an index for the set of masked documents related to the target attribute.

[0081] In Example 2, when a search condition related to a target attribute is given, the search unit 302 uses the index created by the index creation unit 301 to search for masked documents that satisfy the search condition from the set of masked documents.

[0082] In Example 2, the document creation unit 303 creates a target document using the masked document searched by the search unit 302 and the setting string to be set at the masked position of the masked document. That is, the document creation unit 303 creates a target document by setting the setting string corresponding to the masked position for the masked position of the masked document.

[0083] In this way, the document creation processing unit 102 in Example 2 creates a target document from the masked document. In this case, in step S201 of FIG. 12, the index creation unit 301 creates an index for the set of masked documents related to the target attribute, in step S202, the search unit 302 uses the index to search for masked documents that satisfy the search condition from the set of masked documents, and in step S203, the document creation unit 303 creates a target document from the masked document and the setting string. Other points are the same as in Example 1. Note that when creating a target document according to this embodiment, since a template is not required, it is not always necessary to execute the template creation phase. In other words, this embodiment can be implemented as long as there is a set of masked documents even before the execution of the template creation phase.

[0084] · Example 3 Next, Example 3 of the document creation phase will be described. In Example 3 of the document creation phase, the differences from Example 1 will be described, and the description of the components that may be the same as in Example 1 will be omitted.

[0085] <Functional Configuration Example of Document Creation Processing Unit 102> A functional configuration example of the document creation processing unit 102 in Embodiment 3 is shown in FIG. 14. As shown in FIG. 14, the document creation processing unit 102 in Embodiment 3 includes a search unit 302 and a document creation unit 303, and the function of the search unit 302 is different from that in Embodiment 1.

[0086] When a search condition regarding an attribute to be searched (hereinafter, in this embodiment, this attribute will be referred to as the "target attribute") is given, the search unit 302 in Embodiment 3 searches for a template that satisfies the search condition from the template set using a decision tree (overall decision tree, partial decision tree). That is, the search unit 302 searches for a cluster ID that satisfies the search condition regarding the target attribute using a decision tree (overall decision tree, partial decision tree), and acquires a template corresponding to the cluster ID (that is, a template created from a cluster or leaf having that cluster ID).

[0087] In this way, the document creation processing unit 102 in Embodiment 3 searches for a template without creating an index (in other words, using the decision tree as an index). As a result, the index creation unit 301 is not required, and a simpler implementation becomes possible. In this case, step S201 in FIG. 12 is unnecessary, and in step S202, the search unit 302 searches for a template that satisfies the search condition from the template set using a decision tree. Other points are the same as in Embodiment 1.

[0088] [Evaluation Phase] Hereinafter, the evaluation phase will be described.

[0089] <Functional Configuration Example of Evaluation Processing Unit 103> A functional configuration example of the evaluation processing unit 103 in one embodiment is shown in FIG. 15. As shown in FIG. 15, the evaluation processing unit 103 in this embodiment includes a model learning unit 401 and an evaluation unit 402.

[0090] The model learning unit 401 creates a machine learning model (hereinafter also referred to as an evaluation model) for evaluating the clustering results by the cluster creation unit 203. As the evaluation model, any machine learning model that takes as input the document vectors corresponding to each masked document included in the masked document set and the clustering results of those masked documents, and outputs some evaluation score or the like regarding the clustering results can be adopted. The evaluation model is created, for example, by a method such as supervised learning using the document vectors corresponding to each masked document included in the masked document set given as learning data, the clustering results of those masked documents, and the teacher data representing the results of manually classifying those masked documents.

[0091] The evaluation unit 402 uses the document vectors corresponding to each masked document included in the masked document set and the clustering results of those masked documents to calculate an evaluation score by the evaluation model and evaluate the clustering results. For example, when a higher evaluation score indicates better, the evaluation unit 402 evaluates that the clustering result is valid if the evaluation score is equal to or higher than a predetermined threshold, and evaluates that the clustering result is not valid otherwise.

[0092] However, the above evaluation model is just an example and is not limited thereto. For example, instead of the document vector (masked document vector) corresponding to the masked document, a vector corresponding to the document may be input. Also, for example, instead of the document vector, a document ID that identifies the masked document (or the document) corresponding to the document vector may be input. Or, in order to reduce the input dimension of the evaluation model, for example, the statistical value of each document vector (for example, the average value of each document vector, etc.) may be input, or the statistical value of each document ID may be input.

[0093] <Evaluation Process> The evaluation process in this embodiment will be described below with reference to FIG. 16. Here, step S301 in FIG. 16 is performed in advance before evaluating the clustering result, and step S302 is performed each time the clustering result is evaluated. However, when recreating the evaluation model, step S301 is performed at an appropriate timing.

[0094] Step S301: The model learning unit 401 creates an evaluation model by means of supervised learning or the like using the document vectors corresponding to each masked document included in the set of masked documents given as learning data, the clustering results of those masked documents, and the teacher data representing the results of manually classifying those masked documents.

[0095] Step S302: The evaluation unit 402 calculates an evaluation score using the evaluation model with the document vectors corresponding to each masked document included in the set of masked documents and the clustering results of those masked documents, and evaluates the clustering result. Thereby, it is possible to evaluate whether the clustering result by the cluster creation unit 203 is appropriate. For this reason, for example, when the clustering result is not appropriate, it is possible to change the first threshold value and the second threshold value and perform clustering again to create an appropriate template.

[0096] [Supplementary Explanation] Some supplementary matters regarding this embodiment will be described below.

[0097] · The decision tree (overall decision tree, partial decision tree) may be a binary tree or a multi-way tree. It may also be a regression tree or a plurality of decision trees created by a method such as random forest.

[0098] · When creating a decision tree (a global decision tree or a partial decision tree), its depth may be limited. However, if there is bias or the like in the branches of the decision tree and it is non-uniform, the classification accuracy of the decision tree may decrease due to the depth limit. Specifically, there are many cases where document vectors belonging to different clusters are classified into the same leaf. Therefore, in such cases, it is preferable to set a sufficient depth limit or appropriately split the leaves by document vectors belonging to different clusters.

[0099] · In the above embodiment, clustering and decision tree creation were performed using the masked document set, but it is not limited to this. For example, after performing clustering and decision tree creation using a document set, a masked document set may be created. At this time, masked documents for all documents constituting the document set may be created, but for example, only the documents selected as templates may be used as masked documents.

[0100] · In the above evaluation phase, the clustering results were evaluated. However, for example, instead of or in addition to this, a model for evaluating whether the selection of the representative document vector when creating a template is appropriate may be created, and the validity of the representative document vector may be evaluated using this model. In this case, for model learning, document vectors classified into the same cluster and the representative document vector of that cluster are used, and manually selected document vectors are used as teacher data. On the other hand, at the time of evaluation, document vectors classified into the same cluster and the representative document vector of that cluster are used.

[0101] ·When creating a template, it is possible to add or delete certain attributes or add or delete certain sentences. At this time, it is also possible to consider the relationships between attributes or the relationships based on dependencies between sentences. For example, when deleting a certain attribute, it is possible to simultaneously delete the attributes related to that attribute, or when deleting a certain sentence, it is possible to simultaneously delete the sentences related to that sentence in terms of dependency. Similarly, for example, when adding a certain attribute, it is possible to simultaneously add the attributes related to that attribute, or when adding a certain sentence, it is possible to simultaneously add the sentences related to that sentence.

[0102] The present invention is not limited to the specifically disclosed above embodiments, and various modifications, changes, combinations with known technologies, etc. are possible without departing from the description of the claims.

[0103] [References] Reference 1: Extended Named Entities, Internet <URL: http: / / ene-project.info / >

Explanation of Reference Signs

[0104] 10 Document creation device 101 Template creation processing unit 102 Document creation processing unit 103 Evaluation processing unit 104 Storage unit 201 Named entity extraction unit 202 Masking unit 203 Cluster creation unit 204 Decision tree creation unit 204A Overall decision tree creation unit 204B Partial decision tree creation unit 205 End determination unit 206 Template creation unit 301 Index creation unit 302 Search unit 303 Document creation unit 401 Model learning unit 402 Evaluation unit

Claims

1. A cluster creation unit configured to cluster a plurality of document vectors obtained by representing, in vectors, a plurality of documents or a plurality of masked documents in which each part of the plurality of documents is masked, respectively, into clusters of a plurality of granularities; A decision tree creation unit configured to create a decision tree for classifying all or part of the plurality of documents or the plurality of masked documents, with the cluster ID of the cluster of the granularity as the target variable and predetermined information regarding the document as the explanatory variable, in descending order of the granularity; A template creation unit configured to create a template used for creating a target document representing a document that a user desires to create, based on the clustering result or the decision tree; A document creation unit configured to create the target document using a template retrieved from a set of the templates based on a given search condition and the index of the template or the decision tree; An information processing apparatus comprising: The template creation unit: is configured to create, as the template, a masked document corresponding to each document vector selected from each cluster represented by the clustering result or each leaf of the decision tree.

2. The cluster creation unit: is configured to cluster the plurality of document vectors into clusters of a plurality of granularities by a hierarchical clustering method, according to the information processing apparatus of Claim 1.

3. The clusters of the plurality of granularities include at least a first cluster with a coarse granularity and a second cluster with a finer granularity than the first cluster with the coarse granularity. The cluster creation unit: is configured to cluster the plurality of document vectors into the first cluster and cluster the plurality of document vectors into the second cluster, according to the information processing apparatus of Claim 1 or 2.

4. The decision tree creation unit: After creating a first decision tree that classifies the plurality of document vectors with the cluster ID of the first cluster as the target variable and the predetermined information as the explanatory variable, for each of the first clusters, a second decision tree that classifies one or more document vectors belonging to the first cluster with the cluster ID of a second cluster obtained by further clustering the one or more document vectors belonging to the first cluster as the target variable and the predetermined information as the explanatory variable is created. The information processing apparatus according to claim 3, which is configured as such.

5. The information processing apparatus according to claim 1, wherein the predetermined information includes at least one of internal information of the document and external information of the document or information attached to the document.

6. It further has a determination unit configured to determine whether or not to re-perform the clustering by the cluster creation unit and the creation of the decision tree by the decision tree creation unit based on the similarity between the clustering result and the classification result by the decision tree. The cluster creation unit The information processing apparatus according to claim 1, wherein when it is determined by the determination unit to re-perform the clustering, the cluster creation unit is configured to change each threshold corresponding to each of the plurality of granularities and then cluster the plurality of document vectors into clusters of the plurality of granularities.

7. The information processing apparatus according to claim 1, further having an index creation unit configured to create an index of a set of templates created by the template creation unit.

8. The document creation unit The information processing apparatus according to claim 1, wherein the document creation unit is configured to create the target document by setting a given string for a masked portion of the masked document represented by the retrieved template.

9. The information processing apparatus according to claim 1, further having an evaluation unit configured to evaluate the validity of the clustering result by a machine learning model created in advance.

10. A cluster creation procedure for clustering a plurality of document vectors obtained by representing, as vectors, a plurality of documents or a plurality of masked documents in which each part of the plurality of documents is masked into clusters of a plurality of granularities, A decision tree creation procedure for creating a decision tree that classifies all or part of the plurality of documents or the plurality of masked documents, with the cluster ID of the cluster of the granularity as the target variable and predetermined information regarding the document as the explanatory variable, in descending order of the magnitude of the granularity; A template creation procedure for creating a template used for creating a target document representing a document that a user desires to create, based on the clustering result or the decision tree; A document creation procedure for creating the target document using a template retrieved from a set of the templates based on a given search condition and an index of the template or the decision tree; which is executed by a computer; The template creation procedure: An information processing method for creating, as the template, a masked document corresponding to each document vector selected from each cluster represented by the clustering result or each leaf of the decision tree.

11. A cluster creation procedure for clustering a plurality of document vectors each representing a plurality of masked documents in which a plurality of documents or a part of each of the plurality of documents are respectively masked, into clusters of a plurality of granularities; A decision tree creation procedure for creating a decision tree that classifies all or part of the plurality of documents or the plurality of masked documents, with the cluster ID of the cluster of the granularity as the target variable and predetermined information regarding the document as the explanatory variable, in descending order of the magnitude of the granularity; A template creation procedure for creating a template used for creating a target document representing a document that a user desires to create, based on the clustering result or the decision tree; A document creation procedure for creating the target document using a template retrieved from a set of the templates based on a given search condition and an index of the template or the decision tree; which is executed by a computer; The template creation procedure: A program for creating, as the template, a masked document corresponding to each document vector selected from each cluster represented by the clustering result or each leaf of the decision tree.

Citation Information

Patent Citations

  • Document classification apparatus using stereotyped expression, method, program

    JP2005115628A

  • Processing method for time-series analysis of keyword, processing system and computer program thereof

    JP2011141801A