Document Management System

The document management system addresses the issue of incorrect block associations in hierarchical documents by calculating combined similarities across multiple levels, enhancing the accuracy of block correspondence.

JP7865215B2Active Publication Date: 2026-05-26TOYOTA JIDOSHA KK

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
TOYOTA JIDOSHA KK
Filing Date
2023-01-10
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In documents with hierarchical structures, lower hierarchy sentences are often short and similar, leading to incorrect associations between blocks based solely on similarity, as unrelated blocks may be mistakenly linked.

Method used

A document management system that calculates similarities between first-level and higher-level blocks, using a combined similarity metric that increases with both first and second-level similarities, to improve block correspondence accuracy.

Benefits of technology

Enhances the accuracy of block associations by considering both first and second-level similarities, reducing incorrect mappings and improving the precision of document comparisons.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007865215000001
    Figure 0007865215000001
  • Figure 0007865215000002
    Figure 0007865215000002
  • Figure 0007865215000003
    Figure 0007865215000003
Patent Text Reader

Abstract

To provide a document management system that improves accuracy of association between a first-level block in a first document and a first-level block in a second document.SOLUTION: A processing circuit of a document management system calculates a first similarity X1 between a block at a first-level block in a first document and a block at a first level-block in a second document (M12), calculates a similarity between a second-level block in the first document and a second-level block in the second document as a second similarity X2 (M13), calculates, on the basis of the first similarity X1 and the second similarity X2, a composite similarity XA (M15), and associates, on the basis of the composite similarity XA, the first-level block in the first document with the first-level block in the second document (M16).SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a document management system.

Background Art

[0002] Patent Document 1 discloses an example of a document management system that compares a first document and a second document. The documents managed by the system are provided with headings such as chapters and sections and have a hierarchical structure. The system divides the document into a plurality of blocks by classifying the sentences constituting the document for each heading. Then, the system detects the correspondence between the blocks of the first document and the blocks of the second document. Specifically, the system associates a block having a high similarity to the first block of the first sentence among the plurality of blocks of the second document with the first block of the first sentence.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In a document having a hierarchical structure, the sentences in the lower hierarchy are likely to be short sentences, and the same-described sentences frequently appear in the lower hierarchy. Therefore, when the above-described association is made based only on the similarity between the blocks in the lower hierarchy, there is a possibility that a block that is actually unrelated to the first block among the plurality of blocks of the second document may be associated with the first block of the first document.

Means for Solving the Problems

[0005] A document management system for solving the above problems is a system for managing multiple documents, each document being composed of multiple blocks and having a hierarchical structure. The document management system includes a processing circuit. The processing circuit calculates a first similarity, which is the similarity between a first-level block in a first document and a first-level block in a second document. It also calculates a second similarity between a second-level block in the first document that is higher than the first level and a second-level block in the second document, and calculates a combined similarity such that the value increases as the first similarity increases and as the second similarity increases. Based on the combined similarity, the system associates the first-level block in the first document with the first-level block in the second document.

[0006] The document management system described above considers not only the similarity between first-level blocks but also the similarity between second-level blocks to establish a correspondence between first-level blocks in the first document and first-level blocks in the second document. This improves the accuracy of the correspondence between first-level blocks in the first document and first-level blocks in the second document. [Brief explanation of the drawing]

[0007] [Figure 1] Figure 1 is a block diagram illustrating the schematic of the document management system according to the embodiment. [Figure 2] Figure 2 shows examples of the first and second documents managed by the document management system shown in Figure 1. [Figure 3] Figure 3 is a block diagram showing the multiple processes performed by the information processing device of the document management system shown in Figure 1. [Figure 4] Figure 4 is a flowchart showing the mapping process performed by the information processing device shown in Figure 3. [Figure 5]Figure 5 shows the similarity between blocks in the first document and blocks in the second document. [Modes for carrying out the invention]

[0008] An embodiment of a document management system will be described below with reference to Figures 1 to 5. As shown in Figure 1, the document management system 10 is a system for managing multiple documents. The document management system 10 comprises an information processing device 20 and a terminal device 30. The terminal device 30 is configured to send and receive various types of information with the information processing device 20 via a communication line 11 such as a LAN (Local Area Network).

[0009] <Documents under management> Referring to Figure 2, an example of a document managed by the document management system 10 will be explained. For example, this document is a legal document. Figure 2 shows the first document 100A and the second document 100B. For example, the first document 100A is a legal document of country A, and the second document 100B is a legal document of country B that corresponds to the first document 100A.

[0010] Each of the first document 100A and the second document 100B consists of multiple sentences. Both the first document 100A and the second document 100B are hierarchical documents with multiple headings such as chapters, sections, and subsections. In both the first document 100A and the second document 100B, headings are numbered.

[0011] In Document 100A and Document 200B, the heading of Chapter N is denoted by the chapter number as the heading number. The heading of Section M in Chapter N is denoted by the chapter number and section number as the heading number. The heading of Item L in Section M of Chapter N is denoted by the chapter number, section number, and item number as the heading number. Each of "N", "M", and "L" is an integer greater than or equal to 1. For example, the heading of Chapter 3 is denoted by "3." as the heading number. For example, the heading of Section 1 in Chapter 3 is denoted by "3.1." as the heading number. For example, the heading of Item 1 in Section 1 in Chapter 3 is denoted by "3.1.1." as the heading number.

[0012] Each of the first document 100A and the second document 100B can be divided into multiple blocks by separating the text according to headings. For example, the first document 100A has block BK1(1) indicating the heading of Chapter 1, block BK1(2) indicating the heading of Chapter 2, block BK1(3) indicating the heading of Chapter 3, and block BK1(4) indicating the heading of Chapter 4. The first document 100A has block BK1(3.1) indicating the heading of Section 1 of Chapter 3, and block BK1(4.1) indicating the heading of Section 1 of Chapter 4. The first document 100A has block BK1(3.1.1) indicating the heading of Item 1 of Section 1 of Chapter 3, and block BK1(4.1.1) indicating the heading of Item 1 of Section 1 of Chapter 4.

[0013] For example, Document 2 100B has a block BK2(1) indicating the heading of Chapter 1, a block BK2(2) indicating the heading of Chapter 2, a block BK2(3) indicating the heading of Chapter 3, and a block BK2(4) indicating the heading of Chapter 4. Document 2 100B has a block BK2(3.1) indicating the heading of Section 1 of Chapter 3, a block BK2(3.2) indicating the heading of Section 2 of Chapter 3, and a block BK2(4.1) indicating the heading of Section 1 of Chapter 4. Document 2 100B has a block BK2(3.1.1) indicating the heading of Paragraph 1 of Section 1 of Chapter 3, a block BK2(3.2.1) indicating the heading of Paragraph 1 of Section 2 of Chapter 3, and a block BK2(4.1.1) indicating the heading of Paragraph 1 of Section 1 of Chapter 4.

[0014] Furthermore, if the block indicating the heading of Chapter N, Section M, Item L is considered a first-level block, then the block indicating the heading of Chapter N, Section M corresponds to a second-level block, and the block indicating the heading of Chapter N corresponds to a third-level block. For example, in Document 100A, the higher-level block corresponding to block BK1(3.1.1) indicating the heading of Chapter 3, Section 1, Item 1 is block BK1(3.1). Also, the higher-level block corresponding to block BK1(3.1) is block BK1(3) indicating the heading of Chapter 3. For example, in Document 200B, the higher-level block corresponding to block BK2(3.1.1) indicating the heading of Chapter 3, Section 1, Item 1 is block BK2(3.1). Furthermore, the higher-level block corresponding to block BK2(3.1) is block BK2(3), which indicates the heading of Chapter 3. Hereafter, the second block in the higher-level hierarchy corresponding to the first block, which is the first-level block, will be referred to as the "parent-level block" of the first block. The higher-level block corresponding to this second block will be referred to as the "grand-grand-grand-level block" of the first block.

[0015] <Information Processing Device> As shown in Figure 1, the information processing device 20 has a processing circuit 21. The processing circuit 21 is, for example, a microcomputer. In this case, the processing circuit 21 has a CPU 22, a first memory 23, and a second memory 24. The first memory 23 stores a control program executed by the CPU 22. The second memory 24 stores multiple documents managed by the information processing device 20, for example, a first document 100A and a second document 100B. The CPU 22 executes the control program, and the processing circuit 21 analyzes the documents. Specifically, the processing circuit 21 divides the multiple documents 100A and 100B into multiple blocks according to their headings. Then, the processing circuit 21 associates block BK1 of the first document 100A with block BK2 of the second document 100B.

[0016] <Terminal device> The terminal device 30 has a display unit 31 and an operation unit 32 as user interfaces. The display unit 31 displays information transmitted from the information processing device 20, such as the analysis result of a document. The operation unit 32 is operated by an operator based on the information displayed on the display unit 31. The operation unit 32 includes, for example, a keyboard and a mouse. The terminal device 30 transmits information corresponding to the operation of the operation unit 32 by the operator to the information processing device 20.

[0017] <Block Association Processing> Referring to FIGS. 3 to 5, the block association processing M10 executed in the document management system 10 will be described. The block association processing M10 is a series of processes for associating blocks BK1 of the first document 100A and blocks BK2 of the second document 100B after dividing each of the plurality of documents 100A and 100B into a plurality of blocks. The block association processing M10 is executed by the CPU 22 executing the control program of the first memory 23.

[0018] As shown in FIG. 3, the block association processing M10 includes a blocking processing M11, a first similarity calculation processing M12, a second similarity calculation processing M13, a third similarity calculation processing M14, a combined similarity calculation processing M15, and an association processing M16.

[0019] <Blocking Processing> In the blocking processing M11, the processing circuit 21 divides the first document 100A into a plurality of blocks BK1 by classifying the first document 100A by headings as shown in FIG. 2. At this time, the processing circuit 21 searches for the heading numbers from the first document 100A. Then, based on the found heading numbers, the processing circuit 21 divides the first document 100A into a plurality of blocks BK1. Similarly, the processing circuit 21 divides the second document 100B into a plurality of blocks BK2 by classifying the second document 100B by headings.

[0020] <First Similarity Calculation Processing> In the first similarity calculation process M12, the processing circuit 21 calculates a first similarity X1, which is the similarity between a block and another block. For example, when calculating the first similarity X1 between the first block and the second block, the processing circuit 21 calculates the first similarity X1 between the first block and the second block by vectorizing the text of the first block and the text of the second block.

[0021] Note that the similarity between texts can be calculated using, for example, a trained model that has undergone machine learning. In this case, the processing circuit 21 vectorizes the text of the target block. Then, the processing circuit 21 can obtain the similarity between the two blocks as a numerical value by inputting the vectorized text into the trained model. For example, when the text of the first block to be compared and the text of the second block are exactly the same, the similarity is higher than when the text of the first block and the text of the second block do not match. Also, even if the text of the first block and the text of the second block do not exactly match, when the text of the first block and the text of the second block partially match, the similarity is higher than when the text of the first block and the text of the second block do not match at all. Furthermore, when the text of the first block and the text of the second block do not match at all, if the number of characters in the text of the first block and the text of the second block is the same, the similarity is higher than when the number of characters in the text of the first block and the text of the second block is different.

[0022] Examples of the trained model that can calculate the similarity between blocks include, for example, "SentenceBERT". The processing circuit 21 may use a trained model other than "SentenceBERT" as long as it can calculate the similarity between texts.

[0023] In this embodiment, the processing circuit 21 calculates a first similarity X1 between the block BK1 of the first document 100A and the block BK2 of the second document 100B. That is, the processing circuit 21 calculates the first similarity X1 between the block BK1 of the first layer of the first document 100A and the block BK2 of the first layer of the second document 100B.

[0024] Figure 5 illustrates an example of calculating the first similarity X1 between block BK1 (3.1.1) of document 100A and block BK2 of document 200B. Note that the numerical values ​​indicating similarity shown in Figure 5 are just examples. As shown in Figure 2, the text in block BK1 (3.1.1) of document 100A and block BK2 (3.1.1) of document 200B are exactly the same. The text in block BK1 (3.1.1) of document 100A and block BK2 (3.2.1) of document 200B are also exactly the same. The text in block BK1 (3.1.1) of document 100A and block BK2 (4.1.1) of document 200B are also exactly the same. Therefore, the processing circuit 21 calculates "1.00" as the first similarity X1 between block BK1(3.1.1) and block BK2(3.1.1). The processing circuit 21 calculates "1.00" as the first similarity X1 between block BK1(3.1.1) and block BK2(3.2.1). The processing circuit 21 calculates "1.00" as the first similarity X1 between block BK1(3.1.1) and block BK2(4.1.1).

[0025] The processing circuit 21 calculates the first similarity X1 between the first-level block BK1 of the first document 100A and the second-level block BK2 of the second document 100B. As shown in Figure 2, the text in block BK1 (3.1.1) of the first document 100A and block BK2 (3.1) of the second document 100B are different. The text in block BK1 (3.1.1) of the first document 100A and block BK2 (4.1) of the second document 100B are also different. Therefore, as shown in Figure 5, the processing circuit 21 calculates "0.10" as the first similarity X1 between block BK1 (3.1.1) and block BK2 (3.1). The processing circuit 21 calculates "0.10" as the first similarity X1 between block BK1 (3.1.1) and block BK2 (4.1).

[0026] The processing circuit 21 calculates the first similarity X1 between the first-level block BK1 of the first document 100A and the third-level block BK2 of the second document 100B. As shown in Figure 2, the text in block BK1 (3.1.1) of the first document 100A and block BK2 (1) of the second document 100B are different. The text in block BK1 (3.1.1) of the first document 100A and block BK2 (2) of the second document 100B are different in both the number of characters and the content. The text in block BK1 (3.1.1) of the first document 100A and block BK2 (3) of the second document 100B are different. The text in block BK1 (3.1.1) of the first document 100A and block BK2 (4) of the second document 100B are different. Therefore, as shown in Figure 5, the processing circuit 21 calculates "0.10" as the first similarity X1 between block BK1(3.1.1) and block BK2(1). The processing circuit 21 calculates "0.05" as the first similarity X1 between block BK1(3.1.1) and block BK2(2). The processing circuit 21 calculates "0.10" as the first similarity X1 between block BK1(3.1.1) and block BK2(3). The processing circuit 21 calculates "0.10" as the first similarity X1 between block BK1(3.1.1) and block BK2(4).

[0027] The processing circuit 21 also calculates the first similarity X1 between the multiple blocks BK1 in the first document 100A, excluding block BK1(3.1.1), and the multiple blocks BK2 in the second document 100B.

[0028] Furthermore, the processing circuit 21 also calculates the first similarity X1 between blocks BK1 of the first document 100A. <Second Similarity Calculation Process> In the second similarity calculation process M13, the processing circuit 21 calculates the second similarity X2, which is the similarity between a block and other blocks. For example, when calculating the second similarity X2 between the first block and the second block, the processing circuit 21 vectorizes the text of the parent block of the first block and the text of the parent block of the second block. Then, the processing circuit 21 calculates the similarity between the parent block of the first block and the parent block of the second block as the second similarity X2 between the first block and the second block. The processing circuit 21 uses the trained model described above to calculate the similarity between the parent block of the first block and the parent block of the second block, similar to the first similarity calculation process M12 described above.

[0029] In this embodiment, the processing circuit 21 calculates the similarity between the second-level block BK1 of the first document 100A and the second-level block BK2 of the second document 100B as the second similarity X2 between the first-level block BK1 of the first document 100A and the first-level block BK2 of the second document 100B.

[0030] Figure 5 illustrates an example of calculating the second similarity X2 between block BK1(3.1.1) in the first document 100A and block BK2 in the second document 100B. The processing circuit 21 calculates the similarity between block BK1(3.1) in the second layer of the first document 100A and block BK2(3.1) in the second layer of the second document 100B as the second similarity X2 between block BK1(3.1.1) and block BK2(3.1.1). Block BK1(3.1) is the parent layer block of block BK1(3.1.1). Block BK2(3.1) is the parent layer block of block BK2(3.1.1). As shown in Figure 2, the text in block BK1(3.1) of the first document 100A and block BK2(3.1) of the second document 100B are identical. Therefore, the processing circuit 21 calculates "1.00" as the second similarity X2 between block BK1(3.1.1) and block BK2(3.1.1).

[0031] The processing circuit 21 calculates the similarity between block BK1(3.1) in the second level of the first document 100A and block BK2(3.2) in the second level of the second document 100B as the second similarity X2 between block BK1(3.1.1) and block BK2(3.2.1). Block BK1(3.1) is the parent block of block BK1(3.1.1). Block BK2(3.2) is the parent block of block BK2(3.2.1). As shown in Figure 2, the text in block BK1(3.1) of the first document 100A and block BK2(3.2) of the second document 100B do not match. Therefore, the processing circuit 21 calculates "0.10" as the second similarity X2 between block BK1(3.1.1) and block BK2(3.2.1).

[0032] The processing circuit 21 calculates the similarity between block BK1(3.1) in the second level of the first document 100A and block BK2(4.1) in the second level of the second document 100B as the second similarity X2 between block BK1(3.1.1) and block BK2(4.1.1). Block BK1(3.1) is the parent block of block BK1(3.1.1). Block BK2(4.1) is the parent block of block BK2(4.1.1). As shown in Figure 2, the text in block BK1(3.1) of the first document 100A and block BK2(4.1) of the second document 100B are a perfect match. Therefore, the processing circuit 21 calculates "0.10" as the second similarity X2 between block BK1(3.1.1) and block BK2(4.1.1).

[0033] In this embodiment, the processing circuit 21 calculates the similarity between the second-level block BK1 of the first document 100A and the third-level block BK2 of the second document 100B as the second similarity X2 between the first-level block BK1 of the first document 100A and the second-level block BK2 of the second document 100B.

[0034] For example, the processing circuit 21 calculates the similarity between block BK1(3.1) in the second level of the first document 100A and block BK2(3) in the third level of the second document 100B as the second similarity X2 between block BK1(3.1.1) and block BK2(3.1). Block BK1(3.1) is the parent block of block BK1(3.1.1). Block BK2(3) is the parent block of block BK2(3.1). As shown in Figure 2, the text in block BK1(3.1) of the first document 100A and block BK2(3) of the second document 100B do not match. Therefore, the processing circuit 21 calculates "0.10" as the second similarity X2 between block BK1(3.1.1) and block BK2(3.1).

[0035] Furthermore, the processing circuit 21 also calculates the second similarity X2 between the multiple blocks BK1 in the first document 100A, excluding block BK1(3.1.1), and the multiple blocks BK2 in the second document 100B.

[0036] Furthermore, the processing circuit 21 also calculates the second similarity X2 between blocks BK1 of the first document 100A. <Third Similarity Calculation Process> In the third similarity calculation process M14, the processing circuit 21 calculates the third similarity X3, which is the similarity between a block and other blocks. For example, when calculating the third similarity X3 between the first block and the second block, the processing circuit 21 vectorizes the text of the parent-parent-level blocks of the first block and the text of the parent-parent-level blocks of the second block. Then, the processing circuit 21 calculates the similarity between the parent-parent-level blocks of the first block and the parent-parent-level blocks of the second block as the third similarity X3 between the first block and the second block. The processing circuit 21 uses the trained model described above to calculate the similarity between the parent-parent-level blocks of the first block and the parent-parent-level blocks of the second block, similar to the first similarity calculation process M12 described above.

[0037] In this embodiment, the processing circuit 21 calculates the similarity between the third-level block BK1 of the first document 100A and the third-level block BK2 of the second document 100B as the third-level similarity X3 between the first-level block BK1 of the first document 100A and the first-level block BK2 of the second document 100B.

[0038] Figure 5 illustrates an example of calculating the third similarity X3 between block BK1(3.1.1) in the first document 100A and block BK2 in the second document 100B. The processing circuit 21 calculates the similarity between block BK1(3) in the third level of the first document 100A and block BK2(3) in the third level of the second document 100B as the third similarity X3 between block BK1(3.1.1) and block BK2(3.1.1). Block BK1(3) is the parent-parent level block of block BK1(3.1.1). Block BK2(3) is the parent-parent level block of block BK2(3.1.1). As shown in Figure 2, although the text in block BK1(3) of the first document 100A and block BK2(3) of the second document 100B are not exactly identical, both contain the word "general". Therefore, the processing circuit 21 calculates "0.50" as the third similarity X3 between block BK1(3.1.1) and block BK2(3.1.1).

[0039] The processing circuit 21 calculates the similarity between the third-level block BK1(3) in the first document 100A and the third-level block BK2(3) in the second document 100B as the third-level similarity X3 between block BK1(3.1.1) and block BK2(3.2.1). In this case, the processing circuit 21 calculates "0.50" as the third-level similarity X3 between block BK1(3.1.1) and block BK2(3.2.1).

[0040] The processing circuit 21 calculates the similarity between block BK1(3) in the third level of the first document 100A and block BK2(4) in the third level of the second document 100B as the third-level similarity X3 between block BK1(3.1.1) and block BK2(4.1.1). Block BK1(3) is the parent-parent level block of block BK1(3.1.1). Block BK2(4) is the parent-parent level block of block BK2(4.1.1). As shown in Figure 2, the text in block BK1(3) of the first document 100A and block BK2(4) of the second document 100B do not match. Therefore, the processing circuit 21 calculates "0.10" as the third-level similarity X3 between block BK1(3.1.1) and block BK2(4.1.1).

[0041] Furthermore, the processing circuit 21 also calculates the third similarity X3 between the multiple blocks BK1 in the first document 100A, excluding block BK1(3.1.1), and the multiple blocks BK2 in the second document 100B.

[0042] Furthermore, the processing circuit 21 also calculates the third similarity X3 between blocks BK1 of the first document 100A. <Composite Similarity Calculation Process> In the combined similarity calculation process M15, the processing circuit 21 calculates the combined similarity XA between block BK1 in the first document 100A and block BK2 in the second document 100B. At this time, the processing circuit 21 calculates the combined similarity XA based on at least the first similarity X1 among the first similarity X1, second similarity X2, and third similarity X3. That is, the processing circuit 21 calculates the combined similarity XA as shown below.

[0043] The processing circuit 21 calculates the composite similarity XA such that the value increases as the first similarity X1 increases. When the processing circuit 21 calculates the composite similarity XA based on the second similarity X2, it calculates the composite similarity XA such that the value increases as the second similarity X2 increases.

[0044] When the processing circuit 21 calculates the composite similarity XA based on the third similarity X3, it calculates the composite similarity XA such that the value increases as the third similarity X3 increases. In this embodiment, the processing circuit 21 calculates the combined similarity XA between the first-level block BK1 in the first document 100A and the first-level block BK2 in the second document 100B, based on the first similarity X1, second similarity X2, and third similarity X3. For example, the processing circuit 21 calculates the combined similarity XA between block BK1(3.1.1) and block BK2(3.1.1) based on the first similarity X1, second similarity X2, and third similarity X3 between block BK1(3.1.1) and block BK2(3.1.1).

[0045] The processing circuit 21 calculates the combined similarity XA between the first-level block BK1 in the first document 100A and the second-level block BK2 in the second document 100B, based on the first similarity X1 and the second similarity X2. For example, the processing circuit 21 calculates the combined similarity XA between block BK1(3.1.1) and block BK2(3.1) based on the first similarity X1 and the second similarity X2 between block BK1(3.1.1) and block BK2(3.1).

[0046] The processing circuit 21 calculates the combined similarity XA between the first-level block BK1 in the first document 100A and the third-level block BK2 in the second document 100B, based on the first similarity X1. For example, the processing circuit 21 calculates the combined similarity XA between block BK1(3.1.1) and block BK2(3), based on the first similarity X1 between block BK1(3.1.1) and block BK2(3).

[0047] Furthermore, the processing circuit 21 also calculates the combined similarity XA between the multiple blocks BK1 in the first document 100A, excluding block BK1(3.1.1), and the multiple blocks BK2 in the second document 100B.

[0048] Furthermore, the processing circuit 21 also calculates the composite similarity XA between blocks BK1 of the first document 100A. Referring to Figure 4, the processing flow in the composite similarity calculation process M15 will be explained.

[0049] In step S11, the processing circuit 21 obtains the first similarity X1 calculated in the first similarity calculation process M12, the second similarity X2 calculated in the second similarity calculation process M13, and the third similarity X3 calculated in the third similarity calculation process M14.

[0050] In step S13, the processing circuit 21 determines whether the second similarity X2 obtained in step S11 is less than or equal to the judgment value X2th. A low second similarity X2 means that the parent block of block BK2 in the second document 100B is not similar to the parent block of block BK1 in the first document 100A, and therefore it can be considered that there is no need to consider the second similarity X2 when inferring the relationship between block BK1 and block BK2. Thus, the judgment value X2th is set as the criterion for determining whether or not it is necessary to consider the second similarity X2 when inferring the relationship between block BK1 and block BK2. If the second similarity X2 is less than or equal to the judgment value X2th (S13: YES), the processing circuit 21 proceeds to step S15. On the other hand, if the second similarity X2 is higher than the judgment value X2th (S13: NO), the processing circuit 21 proceeds to step S17.

[0051] In step S15, the processing circuit 21 sets the second similarity X2 to 0 (zero). Then, the processing circuit 21 proceeds to step S17. In step S17, the processing circuit 21 determines whether the third similarity X3 obtained in step S11 is less than or equal to the judgment value X3th. A low third similarity X3 means that the parent-parent hierarchical block of block BK2 in the second document 100B is not similar to the parent-parent hierarchical block of block BK1 in the first document 100A, and therefore it can be considered that there is no need to consider the third similarity X3 when inferring the relationship between block BK1 and block BK2. Thus, the judgment value X3th is set as the criterion for determining whether or not it is necessary to consider the third similarity X3 when inferring the relationship between block BK1 and block BK2. If the third similarity X3 is less than or equal to the judgment value X3th (S17: YES), the processing circuit 21 proceeds to step S19. On the other hand, if the third similarity X3 is higher than the judgment value X3th (S17: NO), the processing circuit 21 proceeds to step S21.

[0052] In step S19, the processing circuit 21 sets the third similarity X3 to 0 (zero). Then, the processing circuit 21 proceeds to step S21. In step S21, the processing circuit 21 determines whether the first similarity score X1 obtained in step S11 is higher than the judgment value X1th. If the first similarity score X1 is low, it means that block BK2 of the second document 100B is not similar to block BK1 of the first document 100A, and therefore it can be considered that there is no need to consider the first similarity score X1 when inferring the relationship between block BK1 and block BK2. Therefore, the judgment value X1th is set as the criterion for determining whether or not it is necessary to consider the first similarity score X1 when inferring the relationship between block BK1 and block BK2. If the first similarity score X1 is higher than the judgment value X1th (S21: YES), the processing circuit 21 proceeds to step S23. On the other hand, if the first similarity score X1 is less than or equal to the judgment value X1th (S21: NO), the processing circuit 21 proceeds to step S25.

[0053] In step S23, the processing circuit 21 calculates the composite similarity XA using the following relational expression (D1). In relational expression (D1), "α2" is a correction coefficient that reduces the second similarity X2. "α3" is a correction coefficient that reduces the third similarity X3.

[0054] XA = X1 + X2·α2 + X3·α3 (D1) In this embodiment, the correction coefficients α2 and α3 are set to values ​​greater than 0 (zero) and less than 1, respectively. Therefore, the processing circuit 21 can calculate the composite similarity XA such that the degree to which the second similarity X2 is reflected in the composite similarity XA is smaller than the degree to which the first similarity X1 is reflected in the composite similarity XA. Furthermore, the processing circuit 21 can calculate the composite similarity XA such that the degree to which the third similarity X3 is reflected in the composite similarity XA is smaller than the degree to which the first similarity X1 is reflected in the composite similarity XA.

[0055] Furthermore, a smaller value than the correction coefficient α2 is set as the correction coefficient α3. As a result, the processing circuit 21 can calculate the composite similarity XA such that the degree to which the third similarity X3 is reflected in the composite similarity XA is smaller than the degree to which the second similarity X2 is reflected in the composite similarity XA.

[0056] In step S25, the processing circuit 21 calculates the combined similarity XA as the sum of the second similarity X2 and the third similarity X3. The processing circuit 21 terminates the series of processes after calculating the composite similarity XA in step S23 or step S25.

[0057] <Matching process> As shown in Figure 3, in the mapping process M16, the processing circuit 21 maps block BK1 of the first document 100A to block BK2 of the second document 100B based on the combined similarity XA. For example, the processing circuit 21 links the first block of the first document 100A to one of the multiple blocks BK2 of the second document 100B. At this time, the processing circuit 21 links the second block with the highest combined similarity XA among the multiple blocks BK2 to the first block. The processing circuit 21 also links blocks other than the first block among the multiple blocks BK1 to one of the multiple blocks BK2.

[0058] <Operation of this embodiment> The operation of the document management system 10 will be explained with reference to Figures 2 and 5. In the information processing device 20, the first document 100A is divided into multiple blocks BK1 by separating it according to its headings. Similarly, the second document 100B is divided into multiple blocks BK2 by separating it according to its headings.

[0059] Next, the information processing device 20 calculates the first similarity X1 between block BK1 of the first document 100A and block BK2 of the second document 100B. Furthermore, the similarity between the parent-level block of block BK1 of the first document 100A and the parent-level block of block BK2 of the second document 100B is calculated as the second similarity X2 between block BK1 and block BK2. Finally, the similarity between the parent-highest level block of block BK1 of the first document 100A and the parent-highest level block of block BK2 of the second document 100B is calculated as the third similarity X3 between block BK1 and block BK2.

[0060] Then, the information processing device 20 calculates a combined similarity XA between block BK1 and block BK2 based on the first similarity X1, second similarity X2, and third similarity X3. Based on this combined similarity XA, a correspondence is made between block BK1 of the first document 100A and block BK2 of the second document 100B. For example, among multiple blocks BK2, the block with the highest combined similarity XA with block BK1 is linked to block BK1.

[0061] <Effects of this embodiment> (1) When the relationship between the first block of the first document 100A and the second block of the second document 100B is actually high, the similarity between the parent block of the first block and the parent block of the second block tends to be high. Conversely, when the first block of the first document 100A is actually unrelated to the second block of the second document 100B, the similarity between the parent block of the first block and the parent block of the second block tends to be low. Therefore, when the document management system 10 associates block BK1 of the first document 100A with block BK2 of the second document 100B, it considers not only the first similarity X1 between the blocks but also the second similarity X2. This improves the accuracy of the association between the first-level block BK1 in the first document 100A and the first-level block BK2 in the second document 100B.

[0062] For example, the first similarity X1 between block BK1(3.1.1) of the first document 100A and block BK2(3.2.1) of the second document 100B is equal to the first similarity X1 between block BK1(3.1.1) of the first document 100A and block BK2(3.1.1) of the second document 100B. However, the second similarity X2 between block BK1(3.1.1) and block BK2(3.2.1) is lower than the second similarity X2 between block BK1(3.1.1) and block BK2(3.1.1). Therefore, the combined similarity XA between block BK1(3.1.1) and block BK2(3.2.1) is lower than the combined similarity XA between block BK1(3.1.1) and block BK2(3.1.1). As a result, the document management system 10 can prevent incorrect mapping between block BK1 (3.1.1) and block BK2 (3.2.1).

[0063] (2) When the relationship between the first block of the first document 100A and the second block of the second document 100B is actually high, the similarity between the parent-parent hierarchical block of the first block and the parent-parent hierarchical block of the second block tends to be high. Conversely, when the first block of the first document 100A is actually unrelated to the second block of the second document 100B, the similarity between the parent-parent hierarchical block of the first block and the parent-parent hierarchical block of the second block tends to be low. Therefore, when the document management system 10 associates block BK1 of the first document 100A with block BK2 of the second document 100B, a third similarity factor X3 is also taken into consideration. This further improves the accuracy of the association between the first-level block BK1 in the first document 100A and the first-level block BK2 in the second document 100B.

[0064] For example, the first similarity X1 between block BK1(3.1.1) of document 100A and block BK2(4.1.1) of document 200B is equal to the first similarity X1 between block BK1(3.1.1) of document 100A and block BK2(3.1.1) of document 200B. Also, the second similarity X2 between block BK1(3.1.1) and block BK2(4.1.1) is equal to the second similarity X2 between block BK1(3.1.1) and block BK2(3.1.1). However, the third similarity X3 between block BK1(3.1.1) and block BK2(4.1.1) is lower than the third similarity X3 between block BK1(3.1.1) and block BK2(3.1.1). Therefore, the combined similarity XA between block BK1(3.1.1) and block BK2(4.1.1) is lower than the combined similarity XA between block BK1(3.1.1) and block BK2(3.1.1). As a result, the document management system 10 can prevent incorrect mapping between block BK1(3.1.1) and block BK2(4.1.1).

[0065] (3) The composite similarity XA is calculated such that the degree to which the second similarity X2 is reflected in the composite similarity XA is less than the degree to which the first similarity X1 is reflected in the composite similarity XA. This allows the mapping between block BK1 of the first document 100A and block BK2 of the second document 100B to be performed while considering the similarity of blocks between parent hierarchies to some extent, with the strongest consideration given to the similarity of blocks within the same hierarchical level. This improves the accuracy of the mapping between block BK1 of the first document 100A and block BK2 of the second document 100B.

[0066] (4) The composite similarity XA is calculated such that the degree to which the third similarity X3 is reflected in the composite similarity XA is less than the degree to which the first similarity X1 is reflected in the composite similarity XA. This allows the mapping between block BK1 of the first document 100A and block BK2 of the second document 100B to be performed while considering the similarity between blocks in the parent-parent hierarchy to some extent, with the strongest consideration being given to the similarity between blocks in the same hierarchy. This improves the accuracy of the mapping between block BK1 of the first document 100A and block BK2 of the second document 100B.

[0067] (5) If the second similarity X2 is less than or equal to the judgment value X2th, the composite similarity XA is calculated without using the second similarity X2. Therefore, if there is no relationship between the blocks of the parent hierarchy, the correspondence between block BK1 of the first document 100A and block BK2 of the second document 100B can be made without being affected by the parent hierarchy.

[0068] (6) If the third similarity X3 is less than or equal to the judgment value X3th, the composite similarity XA is calculated without using the third similarity X3. Therefore, if there is no relationship between the blocks in the parent-parent hierarchy, the correspondence between block BK1 of the first document 100A and block BK2 of the second document 100B can be made without being affected by the parent-parent hierarchy.

[0069] <Example of changes> The above embodiment can be implemented with the following modifications. The above embodiment and the following modifications can be combined with each other to the extent that they do not contradict each other technically.

[0070] The processing circuit 21 may calculate the composite similarity XA such that the degree to which the third similarity X3 is reflected in the composite similarity XA is about the same as the degree to which the second similarity X2 is reflected in the composite similarity XA. In this case, it is preferable to make the correction coefficient α3 equal to the correction coefficient α2 in relation (D1).

[0071] The processing circuit 21 may calculate the composite similarity XA such that the degree to which the second similarity X2 is reflected in the composite similarity XA is about the same as the degree to which the first similarity X1 is reflected in the composite similarity XA. In this case, it is preferable to set the correction coefficient α2 to 1 in relation (D1).

[0072] The processing circuit 21 may calculate the composite similarity XA such that the degree to which the third similarity X3 is reflected in the composite similarity XA is about the same as the degree to which the first similarity X1 is reflected in the composite similarity XA. In this case, the correction coefficient α3 in relation (D1) is set to 1.

[0073] The processing circuit 21 may increase the first similarity X1 and calculate the composite similarity XA using the increased first similarity X1. In this case, the processing circuit 21 can make the degree to which the second similarity X2 is reflected in the composite similarity XA smaller than the degree to which the first similarity X1 is reflected in the composite similarity XA, even without decreasing the second similarity X2.

[0074] If the processing circuit 21 calculates the composite similarity XA based on the second similarity X2, it does not need to use the third similarity X3 when calculating the composite similarity XA. The processing circuit 21 may calculate the composite similarity XA based on the second similarity X2 even if the second similarity X2 is less than or equal to the judgment value X2th.

[0075] The processing circuit 21 may calculate the composite similarity XA based on the third similarity X3 even if the third similarity X3 is less than or equal to the judgment value X3th. The processing circuit 21 may calculate the composite similarity XA based on the first similarity X1 even if the first similarity X1 is less than or equal to the judgment value X1th.

[0076] The processing circuit 21 may calculate the first similarity X1, the second similarity X2, and the third similarity X3 using a method different from the method described in the above embodiment. The documents covered may be other documents besides legal documents. Examples of other documents include product instruction manuals and specifications, legal documents, and academic papers.

[0077] The document management system 10 may also be a device for comparing the original document with the revised document. The documents in question do not have to be horizontally written documents like the one shown in Figure 2; they can also be vertically written documents.

[0078] The processing circuit 21 is not limited to one that includes a CPU and ROM and executes software processing. In other words, the processing circuit 21 may have any of the following configurations: (a), (b), or (c).

[0079] (a) The processing circuit 21 comprises one or more processors that perform various processes according to a computer program. The processor includes a CPU and memory such as RAM and ROM. The memory stores program code or instructions configured to cause the CPU to perform the processes. The memory, i.e., computer-readable media, includes any available media that can be accessed by a general-purpose or dedicated computer.

[0080] (b) The processing circuit 21 includes one or more dedicated hardware circuits that perform various processes. Examples of dedicated hardware circuits include application-specific integrated circuits, i.e., ASICs or FPGAs. ASIC is an abbreviation for "Application Specific Integrated Circuit," and FPGA is an abbreviation for "Field Programmable Gate Array."

[0081] (c) The processing circuit 21 includes a processor that executes a portion of the various processes according to a computer program, and a dedicated hardware circuit that executes the remaining processes among the various processes.

[0082] In this specification, the expression "at least one" means "one or more" of the desired options. For example, if there are two options, the expression "at least one" means "only one option" or "both of the two options." As another example, if there are three or more options, the expression "at least one" means "only one option" or "a combination of two or more arbitrary options." [Explanation of Symbols]

[0083] 10...Document management system, 20...Information processing device, 21...Processing circuit, 100A...First document, 100B...Second document, BK1, BK2...Blocks.

Claims

1. A document management system for managing multiple documents that are composed of multiple blocks and have a hierarchical structure, Equipped with a processing circuit, The aforementioned processing circuit is A first similarity score is calculated, which is the similarity between the first-level block in the first document among the multiple documents and the first-level block in the second document among the multiple documents. The similarity between a block in the second level above the first level in the first document and a block in the second level in the second document is calculated as the second similarity between the block in the first level in the first document and the block in the first level in the second document. The composite similarity is calculated such that the value increases as the first similarity increases, and the value increases as the second similarity increases. Based on the composite similarity, a correspondence is made between the first-level blocks in the first document and the first-level blocks in the second document. When calculating the composite similarity, the composite similarity is calculated such that the degree to which the second similarity is reflected in the composite similarity is less than the degree to which the first similarity is reflected in the composite similarity, and when the second similarity is less than or equal to the judgment value, the composite similarity is calculated based only on the first similarity among the first and second similarities. Document management system.

2. The aforementioned processing circuit is The similarity between a block in the third level above the second level in the first document and a block in the third level in the second document is calculated as the third similarity between a block in the first level in the first document and a block in the first level in the second document. When calculating the composite similarity, the composite similarity is calculated such that the value increases as the third similarity increases. The document management system according to claim 1.

3. When the processing circuit calculates the composite similarity, it calculates the composite similarity such that the degree to which the second similarity is reflected in the composite similarity is less than the degree to which the first similarity is reflected in the composite similarity, and the degree to which the third similarity is reflected in the composite similarity is less than the degree to which the second similarity is reflected in the composite similarity. The document management system according to claim 2.