Document processing method, document processing program, and document processing system

The document processing system addresses the challenge of finding similar documents by using an inverted index and similarity calculation, enabling efficient and accurate retrieval of documents with high content similarity.

WO2025143045A1PCT designated stage expired Publication Date: 2025-07-03LEGALON TECHNOLOGIES INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/045973
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-26
Filing Date
2024-12-25
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing document processing systems struggle to efficiently search for documents with similar contents, particularly in managing and searching legal documents like contracts, due to limitations in keyword-based searches.

Method used

A document processing system that utilizes an inverted index database and similarity calculation units to identify and extract documents with high content similarity to a query document, employing filtering processes to reduce computational load.

Benefits of technology

Facilitates efficient retrieval of documents with similar content, reducing computational time and enhancing the accuracy of document management systems, especially for legal documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024045973_03072025_PF_FP_ABST
    Figure JP2024045973_03072025_PF_FP_ABST
Patent Text Reader

Abstract

In this invention, a processor acquires a query document that is a document serving as a search query, calculates, for each of a plurality of search target documents that are documents to be searched on the basis of the query document, the degree of similarity between content of the query document and content of the search target document, extracts a search target document having a high degree of similarity to the query document from among the plurality of search target documents on the basis of the degree of similarity of each of the plurality of search target documents, and outputs information related to the extracted search target document having a high degree of similarity to the query document.
Need to check novelty before this filing date? Find Prior Art

Description

Document processing method, document processing program, document processing system CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application is based on and claims the benefit of priority from Japanese Patent Application No. 2023-218885, filed on December 26, 2023, the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates to a document processing method, a document processing program, and a document processing system.

[0003] A conventional document processing device is described in Patent Document 1. This document processing device displays search keywords based on past search history along with the number of search results for documents searched for using the search keywords.

[0004] Japanese Patent Application Laid-Open No. 2022-114897

[0005] There is a demand for searching for documents with similar contents when it comes to documents such as contracts, etc. The document processing device described in Patent Document 1 has difficulty in meeting such demands.

[0006] An embodiment of the present disclosure provides a document processing method, a document processing program, and a document processing system that are capable of searching for documents with similar content.

[0007] In one embodiment of the document processing method, a processor obtains a query document, which is a document that serves as a search query, calculates the similarity of the content of each of a plurality of search target documents, which are documents that are searched based on the query document, to the content of the query document, extracts search target documents that have a high similarity to the query document from the plurality of search target documents based on the similarity of each of the plurality of search target documents, and outputs information related to the extracted search target documents that have a high similarity to the query document.

[0008] This method makes it possible to search for documents with similar content.

[0009] FIG. 1 is a diagram schematically illustrating an overview of a document processing system according to an embodiment. FIG. 2 is a block diagram illustrating a schematic configuration of the document processing system according to an embodiment. FIG. 3 is a block diagram illustrating a schematic configuration of a document management server device according to an embodiment. FIG. 4 is a diagram schematically illustrating information stored in a document database according to an embodiment. FIG. 5 is a diagram schematically illustrating information stored in an inverted index database according to an embodiment. FIG. 6 is a diagram schematically illustrating an overview of a filtering process according to an embodiment. FIG. 7 is a diagram schematically illustrating a document search method using inverted index information according to an embodiment. FIG. 8 is a sequence chart illustrating an example of the operation of the document processing system according to an embodiment. FIG. 9 is a block diagram illustrating a hardware configuration of a computer according to an embodiment.

[0010] Hereinafter, an embodiment of a document processing method, a document processing program, and a document processing system will be described with reference to the drawings. To facilitate understanding of the description, the same components in each drawing are denoted by the same reference numerals as much as possible, and duplicate descriptions will be omitted.

[0011] First, an overview of a document processing system according to this embodiment will be described.

[0012] (Overview of Document Processing System) As shown in FIG. 1 , the document processing system 10 of this embodiment is a system that manages data of documents D1, D2, D3, etc. created by users in a database DB. The users are, for example, companies that request document information management in the document processing system 10, or personnel within those companies. Documents are information recorded with the intention of being referenced. Examples of documents include official documents, contracts, minutes, rules, notices, regulations, and bylaws. Documents may be in any language, including Japanese, English, Chinese, and other languages. The format of the information in each document is not particularly limited, and may be, for example, a data file created using specific word processing software, or text data extracted from the data file. In this embodiment, a contract is used as an example of a document.

[0013] In such a document processing system 10, for example, when a user is creating data for a specific document Db, there may be a desire to search for other documents whose text is similar to that of the specific document Db. To meet this desire, the document processing system 10 of this embodiment searches for other documents whose content is similar to that of the specific document Db from among the multiple documents stored in the database DB, and presents the other documents found by the search to the user.

[0014] (Configuration of Document Processing System) Next, the configuration of the document processing system of this embodiment will be described.

[0015] 2, the document processing system 10 of this embodiment includes a document management server device 20 and a terminal device 30. The document management server device 20 and the terminal device 30 are connected to each other via a network line N so as to be able to communicate with each other.

[0016] The document management server device 20 is a server-type information processing device that operates in response to requests from the terminal device 30. Note that the document management server device 20 does not necessarily have to be configured as a single information processing device, but may be configured as a combination of multiple information processing devices operating in cooperation with each other, or may operate using any cloud service. Furthermore, the functions of the document management server device 20 may be realized within the terminal device 30.

[0017] The terminal device 30 is an information processing device such as a personal computer (PC) or a tablet terminal.

[0018] The network line N is a communication network capable of high-speed communication, such as the Internet, an intranet, or a wired or wireless communication network such as a LAN (Local Area Network).

[0019] In the document processing system 10, each user is assigned account information. The account information includes a user ID and a password. This account information is transmitted from the terminal device 30 to the document management server device 20 when logging in to the system of the document management server device 20 and is used for user authentication in the document management server device 20. Once authenticated, the user can use various services provided by the document management server device 20 through a web browser or a predetermined application. For example, the user can upload and store content information of a contract created by the user to the document management server device 20 through a web browser or a predetermined application. The document management server device 20 may also store content information of a contract transmitted from any electronic contracting service. The content information of a contract stored in the document management server device 20 is, for example, content information of a contract after it has been concluded. Furthermore, the user can access the database of the document management server device 20 through a web browser or a predetermined application by operating the terminal device 30 to view contracts created by the user in the past or search for a desired contract from among multiple contracts.

[0020] (Configuration of Document Management Server Apparatus) Next, the configuration of the document management server apparatus of this embodiment will be described.

[0021] 3, the document management server device 20 of this embodiment includes a storage unit 21, a communication unit 22, and a control unit 23. The storage unit 21, the communication unit 22, and the control unit 23 are electrically connected to each other.

[0022] The storage unit 21 of this embodiment stores various programs and various data for operating the document management server device 20. For example, the storage unit 21 stores a document database 210 and an inverted index database 211.

[0023] As shown in FIG. 4 , the document database 210 of this embodiment includes, for example, account information, document ID, sequence ID, order information, status information, and document content information for each user. The document ID is, for example, information that can uniquely identify each document stored in the document database 210. The sequence ID is, for example, information that can identify a sequence. The sequence is, for example, information that indicates a series of revision processes. Multiple documents belonging to the same sequence ID indicate that they belong to the same series of revision processes. The order information is, for example, information that indicates the order of revisions numerically and is sometimes called version information. The document content information is, for example, information that indicates the content of a contract (e.g., text, figures, tables). The document content information is, for example, information created or uploaded by a specific user who created a draft of the document, or by a user different from the specific user. The different user is another user who belongs to the same organization as the specific user, for example, another user who can supply document data.

[0024] In document management, etc., it may be necessary to determine to which sequence a certain document belongs. In particular, with legal documents, draft documents created during interactions with counterparties included in a sequence may contain information about the negotiation process, making it highly desirable to appropriately assign the file to the correct sequence for management. In this case, it may be possible to check the similarity of the content of the document with other documents and determine that the sequence containing other documents determined to be similar is the sequence to which the document should belong. However, when creating documents using a template such as a contract, it is expected that there will be many documents that differ only slightly from documents created using the same template, for example, by only a few characters. When extracting similar documents from multiple candidates with similar content, this embodiment, which uses documents as queries rather than keywords, can be suitably used. Furthermore, legal documents such as such contracts often contain large amounts of document data, and this embodiment allows for efficient extraction of similar documents.

[0025] The inverted index database 211 stores inverted index information such as that shown in FIG. 5 for each user. The inverted index information in this embodiment is information that associates, for each specific word contained in a document registered in the document database 210, the specific word with the ID of the document containing that specific word. By using this inverted index information, for example, documents containing the word "stock" can be identified as documents with IDs "D0001," "D0003," and "D0010." Also, for example, a document containing the word "AAA" can be identified as a document with ID "D0001."

[0026] The communication unit 22 of this embodiment performs various communications with, for example, the terminal device 30. The communication unit 22 acquires, for example, operation information performed by a user on the terminal device 30. The communication unit 22 also transmits, to the terminal device 30, various types of image information to be displayed on the terminal device 30.

[0027] The control unit 23 of this embodiment controls, for example, the document management server device 20. The control unit 23 includes, as functional components realized by executing a program stored in the storage unit 21, an inverted index construction unit 231, a query acquisition unit 232, a similarity calculation unit 233, an extraction unit 234, and an output unit 235.

[0028] (Inverted Index Construction Unit) The inverted index construction unit 231 of this embodiment constructs inverted index information stored in the inverted index database 211. For example, every time a new document is registered in the document database 210 or the contents of a document are changed, the inverted index construction unit 231 reads the content information of the document and constructs inverted index information as shown in Fig. 5 from the content information of the read document. Note that the specific words used in the inverted index information may be predetermined words or words specified by the system administrator of the document management server device 20, etc.

[0029] (Query Acquisition Unit) The query acquisition unit 232 of this embodiment acquires a query document, which is a document that serves as a search query. For example, suppose that a user operates a web browser or a predetermined application on the terminal device 30 to specify, as a search query, one of multiple documents that the user can access and that are stored in the document database 210. In this case, the query acquisition unit 232 acquires content information of the document specified by the user as the search query from the document database 210. For example, in the case of User A shown in FIG. 4, the multiple documents that the user can access are multiple documents associated with User A. Hereinafter, a document acquired as a search query by the query acquisition unit 232 will be referred to as a "query document."

[0030] The query acquisition unit 232 may acquire, as a query document, a predetermined document input by a user operating a web browser or a predetermined application on the terminal device 30 .

[0031] Furthermore, acquisition of a query document by the query acquisition unit 232 is not limited to being accompanied by a user operation, but may be performed automatically by the query acquisition unit 232. For example, when a user displays a specific contract on the terminal device 30 or uploads a contract file, the query acquisition unit 232 may automatically set the contract as a query document and acquire information about the contract.

[0032] (Similarity Calculation Unit) The similarity calculation unit 233 of this embodiment calculates the similarity of the content of a search target document to the content of a query document for each of a plurality of search target documents registered in the document database 210. The similarity calculation unit 233 of this embodiment calculates the Jaccard similarity between the content of each of the search target document and the query document. The Jaccard similarity J(x, y) can be calculated, for example, by the following formula f1.

[0033] In formula f1, "x" is a set of words contained in the document to be searched. For example, if the document to be searched contains the content "AAA Co., Ltd.", the set of words x is expressed in the form "x = {stock, company, AAA, ...}" by dividing the document into words. "y" is a set of words contained in the query document.

[0034] Using the similarity J calculated by the above formula f1, for example, if the similarity J satisfies "J≧t" with respect to a predetermined threshold t, it can be determined that the content of the search target document is similar to the content of the query document. Also, if "J<t" is satisfied, it can be determined that the search target document is not similar to the query document. The threshold t is set to satisfy "0<t≦1."

[0035] However, when the number of search target documents is large, performing the above-described similarity calculation for all of them may require a long calculation time. Meanwhile, the multiple search target documents may include documents that are completely different from the query document in addition to documents similar to the query document, as shown in FIG. 6 . Therefore, in order to shorten the similarity calculation time, the similarity calculation unit 233 performs a filtering process to narrow down the multiple search target documents to those that are estimated to be candidate solutions for the search result, as shown in FIG. 6 , and then calculates the similarity only for the narrowed-down search target documents.

[0036] Next, the procedure of the process performed by the similarity calculation unit 233 will be specifically described. The similarity calculation unit 233 first determines whether the number N of search target documents is equal to or greater than a predetermined number Nth. If the number N of search target documents is less than the predetermined number Nth, the similarity calculation unit 233 extracts, from the multiple search target documents, search target documents that satisfy a first filtering condition as first search target documents. The first search target documents are search target documents that are estimated to be candidate solutions for the search results. The first filtering condition is a condition that all of the following conditions (a1) to (a3) ​​are satisfied. Note that the "length of the document content" refers to the number of types of words contained in the specified document. Furthermore, a prefix refers to a subset of the word set of the specified document, from the first word to a predetermined word. (a1) The difference between the length of the query document content and the length of the search target document content is less than a predetermined length. (a2) At least one word contained in the prefix of the first specified query document matches one word contained in the prefix of the first specified search target document. (a3) The degree of match between the words contained in the prefix of the query document and the words contained in the prefix of the search target document is equal to or greater than a predetermined threshold.

[0037] Here, when the length of the content of a specified search target document is |x| and the length of the content of a query document is |y|, the similarity calculation unit 233 determines that the condition (a1) is satisfied if the length of the content of the specified search target document |x| satisfies the following formula f2 using the above threshold value t.

[0038] On the other hand, when a finite set U is the universal set and its subset is x(⊆U), the i-th to j-th subsets are denoted as x[i..j] (1≦i≦j≦|x|). In this case, the prefix of x can be denoted as x[1..i], and the suffix of x can be denoted as x[j..|x|]. When the prefix of a predetermined search target document is x[1..lx] and the prefix of a query document is y[1..ly], the similarity calculation unit 233 determines that the above condition (a2) is satisfied if the following formula f3 is satisfied:

[0039] In formula f3, "lx" and "ly" can be calculated by the following formulas f4 and f5. Also, "T" in formulas f4 and f5 can be calculated by the following formula f6.

[0040] Furthermore, with a given element u (∈U) included in a finite set U as the boundary, the prefix Px(u) and suffix Sx(u) of set x are defined as in the following equations f7 and f8. In equation f7, "a≦u" indicates that element a (∈U) and element u (∈U) are equal or that element a is smaller than element u. In equation f8, "u<a" indicates that element u is smaller than element a.

[0041] The similarity calculation unit 233 determines that the above condition (a3) ​​is satisfied if the following formula f9 is satisfied, where the prefix and suffix of a predetermined search target document are "Px(u)" and "Sx(u)," respectively, and the prefix and suffix of a query document are "Py(u)" and "Sy(u)," respectively. In formula f9, "O((Px(u)), Py(u))" can be calculated using the following formula f10.

[0042] The similarity calculation unit 233 extracts one or more first search target documents that satisfy all of the above conditions (a1) to (a3), and then performs a process of calculating the similarity of the extracted one or more first search target documents of the solution candidates using the above formula f1.

[0043] In this way, if the search target documents of the solution candidates are narrowed down in advance and then similarity is calculated, the calculation time can be shortened compared to calculating similarity for all search target documents. On the other hand, if the number of search target documents N is, for example, 10,000 or more, the number of accesses to the database increases, and the search may end up being slow due to data access. Therefore, when the number of search target documents is particularly large, the similarity calculation unit 233 of this embodiment further narrows down the search target documents of the solution candidates by using the inverted index information stored in the inverted index database 211.

[0044] Specifically, when the number N of search target documents is equal to or greater than a predetermined number Nth, the similarity calculation unit 233 uses the inverted index information to extract, from the multiple search target documents, search target documents that satisfy a second filtering condition as second search target documents. The second search target documents are search target documents that are estimated to be solution candidates for the search results. The second filtering condition is a condition that a specific word contained in the query document is contained in the search target documents. For example, when a character string from the first character to a predetermined character of a specific document is used as a prefix, the similarity calculation unit 233 extracts some words contained in the prefix of the query document as search words, and then extracts search target documents that include the extracted search words based on the inverted index information. Specifically, as shown in FIG. 7 , when the prefix of the query document includes the words {stock, company, XXX, ...}, the similarity calculation unit 233 extracts "stock" and "XXX" as search words. Then, the similarity calculation unit 233 extracts documents with document IDs "D0001," "D0003," and "D0010" as solution candidates from the word "stock" using the inverted index information shown in Fig. 5. Next, the similarity calculation unit 233 executes a process of calculating the similarity between the extracted single or multiple solution candidates and the search target documents.

[0045] (Extraction Unit) The extraction unit 234 of this embodiment extracts, from among a plurality of search target documents, search target documents that have a high degree of similarity to the query document as search result documents.

[0046] For example, if the number N of search target documents is less than a predetermined number Nth, the similarity calculation unit 233 extracts multiple first search target documents and calculates the similarity of each of the multiple first search target documents. In this case, the extraction unit 234 extracts first search target documents whose similarity J satisfies "J≧t" as search result documents that are highly similar to the query document.

[0047] On the other hand, if the number N of search target documents is equal to or greater than a predetermined number Nth, the similarity calculation unit 233 extracts multiple second search target documents and calculates the similarity of each of the multiple second search target documents. In this case, the extraction unit 234 extracts second search target documents whose similarity J satisfies "J≧t" as search result documents that are highly similar to the query document.

[0048] (Output Unit) The output unit 235 of this embodiment transmits information related to one or more search result documents extracted by the extraction unit 234 to the terminal device 30, thereby displaying the information related to the search result documents on the terminal device 30 via a web browser or a predetermined application. The information related to the search result documents may be the title or content information of the search result documents themselves, or may be the title or content information of other documents related to the search result documents. The other documents may be, for example, documents that have the same sequence ID as the search result document and have the most recent order information. Furthermore, the information related to the search result documents may be the titles and content information of multiple documents that have the same sequence ID as the search result document and correspond to each of the multiple order information, as well as the editing history of those documents.

[0049] (Example of Operation of Document Processing System) Next, an example of operation of the document processing system 10 of this embodiment will be described.

[0050] 8 , in the document processing system 10 of this embodiment, when a user first operates the terminal device 30 to specify a query document (step S20), the query acquisition unit 232 of the document management server device 20 acquires the query document specified by the user (step S30). As a result, the similarity calculation unit 233 of the document management server device 20 calculates the similarity of each of the multiple search target documents registered in the document database 210 to the query document. The extraction unit 234 of the document management server device 20 then extracts, as search result documents, search target documents that have a high similarity to the query document from the multiple search target documents.

[0051] Specifically, the similarity calculation unit 233 determines whether the number N of search target documents is equal to or greater than a predetermined number Nth (step S31). If the number N of search target documents is less than the predetermined number Nth (step S31: NO), the similarity calculation unit 233 extracts, from the plurality of search target documents, a plurality of search target documents that satisfy the first filtering condition as first search target documents (step S32), and calculates the similarity of each of the extracted plurality of first search target documents (step S33). Then, the extraction unit 234 extracts, from the plurality of first search target documents, a first search target document whose similarity J satisfies "J≧t" as a search result document that is highly similar to the query document (step S34).

[0052] On the other hand, if the number N of search target documents is equal to or greater than the predetermined number Nth (step S31: YES), the similarity calculation unit 233 extracts, from the plurality of search target documents, a plurality of search target documents that satisfy the second filtering condition as second search target documents (step S35), and calculates the similarity of each of the extracted plurality of second search target documents (step S36).Then, the extraction unit 234 extracts, from the plurality of second search target documents, second search target documents whose similarity J satisfies "J≧t" as search result documents that are highly similar to the query document (step S37).

[0053] Next, the output unit 235 of the document management server device 20 transmits information related to the search result documents extracted by the processing of step S34 or step S37 to the terminal device 30 (step S38), which causes the terminal device 30 to display the information related to the search result documents (step S21).

[0054] The user then performs a predetermined operation. For example, the user can associate the query document or a document related to the query document with the sequence to which the search result document belongs, according to a suggestion from the system (a type of information related to the search result document). In this case, the query document may be recorded and treated as the document with the latest order information of the documents in the sequence, based on information such as the upload date or update date.

[0055] (Hardware Configuration of Document Processing System) Next, an example of a hardware configuration in which the document management server device 20 and the terminal device 30 are realized by a computer 1000 will be described with reference to Fig. 9. Fig. 9 is a diagram showing an example of the hardware configuration of the computer 1000.

[0056] 9, the computer 1000 includes, for example, a processor 1001, a memory 1002, a storage device 1003, an input I / F unit 1004, a data I / F unit 1005, a communication I / F unit 1006, and a display device 1007. Note that the computer 1000 may include a plurality of each of the processor 1001, the memory 1002, the storage device 1003, the input I / F unit 1004, the data I / F unit 1005, the communication I / F unit 1006, and the display device 1007.

[0057] The computer 1000 may be, for example, a server computer, a personal computer (e.g., desktop, laptop, tablet, etc.), a media computing platform (e.g., cable, satellite set-top box, digital video recorder, etc.), a handheld computing device (e.g., PDA, email client, etc.), or any other type of computing or communications platform.

[0058] The processor 1001 is a control unit that controls various processes in the computer 1000 by, for example, executing a program stored in the memory 1002. Each functional unit of the document management server device 20 and the terminal device 30 can be realized by, for example, the processor 1001 executing a program stored in the storage device 1003, or by a machine learning model stored in the storage device 1003.

[0059] The memory 1002 is a storage medium such as a RAM (Random Access Memory), etc. The memory 1002 temporarily stores the program code of the program executed by the processor 1001 and data required when the program is executed.

[0060] The storage device 1003 is a non-volatile storage medium such as a hard disk drive (HDD), flash memory, etc. The storage device 1003 stores an operating system and various programs for realizing the above-mentioned components.

[0061] The input I / F unit 1004 is a device for receiving input from a user. The input I / F unit 1004 is, for example, a keyboard, a mouse, a touch panel, various sensors, a wearable device, etc. The input I / F unit 1004 may be connected to the computer 1000 via an interface such as a USB (Universal Serial Bus).

[0062] The data I / F unit 1005 is a device for inputting data from outside the computer 1000. The data I / F unit 1005 is, for example, a drive device for reading data stored in various storage media. The data I / F unit 1005 may be provided outside the computer 1000. When the data I / F unit 1005 is provided outside the computer 1000, the data I / F unit 1005 is connected to the computer 1000 via an interface such as a USB.

[0063] The communication I / F unit 1006 is a device for performing data communication via a network such as the Internet, either wired or wirelessly, with devices external to the computer 1000. The communication I / F unit 1006 may be provided external to the computer 1000. When the communication I / F unit 1006 is provided external to the computer 1000, the communication I / F unit 1006 is connected to the computer 1000 via an interface such as a USB.

[0064] The display device 1007 is a device for displaying various types of information. The display device 1007 is, for example, a liquid crystal display, an organic EL (Electro-Luminescence) display, a display of a wearable device, or the like. The display device 1007 may be provided external to the computer 1000. When the display device 1007 is provided external to the computer 1000, the display device 1007 is connected to the computer 1000 via, for example, a display cable. Furthermore, when a touch panel is used as the input I / F unit 1004, the display device 1007 may be configured as an integral part of the input I / F unit 1004.

[0065] (Functions and Effects of the Document Processing System of the Present Embodiment) As described above, the document processing system 10 of the present embodiment includes a query acquisition unit 232, a similarity calculation unit 233, an extraction unit 234, and an output unit 235. The query acquisition unit 232 acquires a query document, which is a document that serves as a search query. The similarity calculation unit 233 calculates the similarity of the content of each of a plurality of search target documents, which are documents to be searched based on the query document, to the content of the query document. The extraction unit 234 extracts, from the plurality of search target documents, search target documents that have a high similarity to the query document, based on the similarity of each of the plurality of search target documents. The output unit 235 outputs information related to the search target documents that have a high similarity to the query document, extracted by the extraction unit 234.

[0066] With this configuration, when a user specifies a query document, search target documents that have a high degree of similarity to the content of the query document are extracted, and information related to the search target documents is output, making it possible to search for documents with similar content.

[0067] When calculating the similarity, if the number N of search target documents is less than a predetermined number Nth, the similarity calculation unit 233 extracts, from the plurality of search target documents, search target documents that satisfy the first filtering condition as first search target documents, and calculates the similarity of the extracted first search target documents. The first filtering condition is a condition that all of the above conditions (a1) to (a3) ​​are satisfied.

[0068] According to this configuration, when the number N of search target documents is less than a predetermined number Nth, the search target documents are narrowed down from the plurality of search target documents to those that satisfy the first filtering condition as first search target documents, and then the similarity of the narrowed down first search target documents is calculated. Therefore, it is possible to reduce the time required to calculate the similarity compared to when the similarity of all search target documents is calculated.

[0069] When calculating the similarity, if the number N of search target documents is equal to or greater than a predetermined number Nth, the similarity calculation unit 233 reads the transposed index information from the memory unit 21, and based on the transposed index information, extracts second search target documents that contain words in the query document from among the multiple search target documents, and calculates the similarity of the second search target documents.

[0070] According to this configuration, when the number N of search target documents is equal to or greater than a predetermined number Nth, the search target documents are narrowed down from the plurality of search target documents to those that satisfy the second filtering condition as second search target documents, and then the similarity of the narrowed down second search target documents is calculated, thereby making it possible to reduce the time required to calculate the similarity compared to when the similarity of all search target documents is calculated.

[0071] Other Embodiments The present disclosure is not limited to the above specific examples.

[0072] For example, documents that the document processing system 10 targets are not limited to contracts, but may also be official documents, minutes, rules, notices, regulations, and rules.

[0073] The first and second filtering conditions can be changed as appropriate. For example, the first filtering condition may be a condition that at least one of the above conditions (a1) to (a3) ​​is satisfied.

[0074] 8, the output unit 235 may extract a predetermined number of first search target documents with the highest similarity J from among the plurality of first search target documents as search result documents with high similarity to the query document. Furthermore, the output unit 235 may perform a similar process on the second search target documents in the process of step S37 shown in FIG.

[0075] The functional configurations of the document management server device 20 and the terminal device 30 may be provided only in the document management server device 20, only in the terminal device 30, or in both the document management server device 20 and the terminal device 30. For example, the inverted index construction unit 231, the query acquisition unit 232, the similarity calculation unit 233, the extraction unit 234, and the output unit 235 of the document management server device 20 may be provided in the terminal device 30. Furthermore, the document processing system 10 may include a plurality of document management server devices 20 and a plurality of terminal devices 30.

[0076] Design modifications made by a person skilled in the art to the above specific examples as appropriate are also included within the scope of the present disclosure as long as they comprise the features of the present disclosure. The elements of each of the above specific examples, as well as their arrangement, conditions, shape, etc., are not limited to those exemplified and can be modified as appropriate. The elements of each of the above specific examples can be combined as appropriate as long as no technical contradictions arise.

Claims

1. A document processing method, wherein a processor obtains a query document which is a document serving as a search query, calculates a similarity between the content of the search target document which is a document to be searched based on the query document and the content of the query document for each of a plurality of search target documents, extracts a search target document having a high similarity to the query document from among the plurality of search target documents based on the similarity of each of the plurality of search target documents, and outputs information related to the extracted search target document having a high similarity to the query document.

2. The document processing method according to claim 1, wherein when the number of the search target documents is less than a predetermined number when the processor calculates the similarity, the processor extracts, as a first search target document, a search target document that satisfies a predetermined filtering condition from among the plurality of search target documents, and calculates the similarity of the first search target document.

3. The document processing method according to claim 2, wherein the processor uses, as the predetermined filtering condition, a condition that a difference between the length of the content of the query document and the length of the content of the search target document is less than a predetermined length.

4. When a subset from the first word to the nth word of the word set of a predetermined document is used as a prefix, the document processing method according to claim 2, wherein the processor uses, as the predetermined filtering condition, a condition that at least one of the words included in the prefix of the query document up to the nth word and the words included in the prefix of the search target document up to the nth word matches.

5. When a subset from the first word to the nth word of the word set of a predetermined document is used as a prefix, the document processing method according to claim 2, wherein the processor uses, as the predetermined filtering condition, a condition that a degree of coincidence between the words included in the prefix of the query document and the words included in the prefix of the search target document is equal to or more than a predetermined threshold value.

6. The method for document processing according to claim 1, wherein when calculating the similarity, if the number of search target documents is equal to or more than a predetermined number, the processor reads from the storage unit inverted index information capable of specifying, from the words included in each of the plurality of search target documents, the search target documents including the word, extracts from the plurality of search target documents a second search target document including the word of the query document based on the inverted index information, and calculates the similarity of the second search target document.

7. The method for document processing according to claim 1, wherein the processor calculates the Jaccard similarity as the similarity.

8. The method for document processing according to claim 1, wherein the processor obtains the plurality of search target documents from a database.

9. The method for document processing according to claim 1, wherein the query document and the search target documents are contract documents.

10. A program for causing a processor to execute a process of obtaining a query document which is a document serving as a search query, a process of calculating, for each of a plurality of search target documents which are documents searched based on the query document, the similarity of the content of the search target document with respect to the content of the query document, a process of extracting, from among the plurality of search target documents, a search target document having a high similarity with the query document based on the similarity of each of the plurality of search target documents, and a process of outputting information related to the search target document having a high similarity with the query document, which has been extracted.

11. A document processing system comprising: a query acquisition unit that acquires a query document which is a document serving as a search query; a similarity calculation unit that calculates, for each of a plurality of search target documents which are documents searched based on the query document, the similarity of the content of the search target document with respect to the content of the query document; an extraction unit that extracts, from among the plurality of search target documents, a search target document having a high similarity with the query document based on the similarity of each of the plurality of search target documents calculated by the similarity calculation unit; and an output unit that outputs information related to the search target document having a high similarity with the query document, which has been extracted by the extraction unit.

Citation Information

Patent Citations

  • Similar data search device, similar data search method, and computer-readable storage medium

    WO2014136810A1