Document processing method, document processing program, and document processing system

The document processing system addresses the challenge of identifying similar documents by using Jaccard similarity and filtering techniques, enhancing search efficiency and accuracy for documents with minor content variations.

JP2025101835APending Publication Date: 2025-07-08LEGALON TECHNOLOGIES INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023218885
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing document processing systems struggle to efficiently search for documents with similar contents, particularly in managing and searching legal documents like contracts, where small differences in content can complicate the identification of related documents.

Method used

A document processing system that calculates the similarity between query documents and target documents using the Jaccard similarity index, employing filtering processes to narrow down search candidates, and utilizes inverted index information to expedite similarity calculations for large document sets.

Benefits of technology

Enables efficient identification of documents with similar contents, reducing calculation time and improving search accuracy by focusing on high-similarity documents, especially in large datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025101835000001_ABST
    Figure 2025101835000001_ABST
Patent Text Reader

Abstract

To provide a document processing system capable of retrieving documents having similar contents.SOLUTION: A document processing system 10 is provided with a query acquisition unit 232, a similarity calculation unit 233, an extraction unit 234, and an output unit 235. A query acquisition unit 232 acquires a query document as a document to be a retrieval query. A similarity calculation unit 233 calculates the similarity of the contents of a retrieval object document to the contents of the query document for each of a plurality of retrieval object documents that is retrieved based on the query document. An extraction unit 234 extracts a retrieval object document having a high similarity with the query document from the plurality of retrieval object documents based on the similarity of each of the plurality of retrieval object documents. An output unit 235 outputs information related to the retrieval object document having a high similarity with the query document extracted by the extraction unit 234.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a document processing method, a document processing program, and a document processing system.

Background Art

[0002] Conventionally, there is a document processing apparatus described in Patent Document 1 below. This document processing apparatus displays a search keyword based on a past search history together with the number of search results of documents searched with the search keyword.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Regarding documents such as contracts, there is a demand to search for documents with similar contents. It is difficult for the document processing apparatus described in Patent Document 1 to meet such a demand.

[0005] One embodiment of the present disclosure aims to provide a document processing method, a document processing program, and a document processing system capable of searching for documents with similar contents.

Means for Solving the Problems

[0006] In a document processing method according to an embodiment, a processor acquires a query document that is a document serving as a search query, calculates a similarity between the content of the query document and the content of each of a plurality of target documents to be searched based on the query document, extracts, from among the plurality of target documents, a target document having a high similarity to the query document based on the similarity of each of the plurality of target documents, and outputs information related to the target document having a high similarity to the query document that has been extracted.

[0007] According to this method, it is possible to search for documents with similar contents.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Modes for Carrying Out the Invention

[0009] Hereinafter, an embodiment of a document processing method, a document processing program, and a document processing system will be described with reference to the drawings. For ease of understanding of the description, the same reference numerals are given to the same components in each drawing as much as possible, and duplicate descriptions are omitted.

[0010] <Embodiment> First, the outline of the document processing system of the present embodiment will be described.

[0011] (Outline of the Document Processing System) As shown in FIG. 1, the document processing system 10 of the present embodiment is a system that manages data of documents D1, D2, D3, ··· created by a user in a database DB. The user is, for example, a company that desires to manage document information in the document processing system 10, or a person in charge of that company. A document is information recorded on the premise of being referred to. Documents include, for example, official documents, contracts, minutes of meetings, rules, circulars, regulations, and bylaws. The language of the document is not limited, and includes those described in Japanese, as well as those described in English, Chinese, etc. The format of the information of each document is not particularly limited, and may be, for example, a data file created by a specific word processing software, or text data extracted from the data file. In the present embodiment, a contract will be described as an example of a document.

[0012] In such a document processing system 10, for example, when a user is creating data of a predetermined document Db, there is a desire to search for other documents whose text is similar to that of the predetermined document Db. In order to meet such a desire, the document processing system 10 of the present embodiment searches for other documents whose content is similar to that of a predetermined document Db from among a plurality of documents stored in the database DB, and presents the other documents obtained by the search to the user.

[0013] (Configuration of Document Processing System) Next, the configuration of the document processing system of the present embodiment will be described.

[0014] As shown in FIG. 2, the document processing system 10 of the present embodiment includes a document management server device 20 and a terminal device 30. The document management server device 20 and the terminal device 30 are communicably connected to each other via a network line N.

[0015] The document management server device 20 is a server-type information processing device and operates in response to requests from the terminal device 30. Note that the document management server device 20 does not necessarily have to be configured as a single information processing device, and it may be configured such that a plurality of information processing devices cooperate to operate, or it may operate by means of an arbitrary cloud service. Further, the functions of the document management server device 20 may be realized within the terminal device 30.

[0016] The terminal device 30 is an information processing device such as a PC (Personal Computer) or a tablet terminal.

[0017] The network line N is a communication network capable of high-speed communication, and is, for example, a wired or wireless communication network such as the Internet, an intranet, or a LAN (Local Area Network).

[0018] In the document processing system 10, account information is assigned to each user. The account information includes a user ID and a password. This account information is transmitted from the terminal device 30 to the document management server device 20 when logging in to the system of the document management server device 20, and is used for user authentication in the document management server device 20. When this authentication is successful, the user can use various services provided by the document management server device 20 through a Web browser or a predetermined application. For example, the user can upload and save the content information of a contract document created by the user to the document management server device 20 through a Web browser or a predetermined application. Further, the document management server device 20 may save the content information of a contract document transmitted from an arbitrary electronic signature service. The content information of the contract document stored in the document management server device 20 is, for example, the content information of the contract document after conclusion. Furthermore, the user can access the database of the document management server device 20 by operating the terminal device 30 through a Web browser or a predetermined application, and can view a contract document created by the user in the past or search for a desired contract document from among a plurality of contract documents.

[0019] (Configuration of the Document Management Server Device) Next, the configuration of the document management server device of the present embodiment will be described.

[0020] As shown in FIG. 3, the document management server device 20 of the present embodiment includes a storage unit 21, a communication unit 22, and a control unit 23. The storage unit 21, the communication unit 22, and the control unit 23 are electrically connected.

[0021] In the storage unit 21 of the present embodiment, various programs and various data for operating the document management server device 20 are stored. For example, a document database 210 and an inverted index database 211 are stored in the storage unit 21.

[0022] As shown in FIG. 4, the document database 210 of the present embodiment includes, for example, for each user, account information, document ID, sequence ID, order information, status information, and content information of the document. The document ID is information that can uniquely identify each document stored in the document database 210, for example. The sequence ID is information that can identify a sequence, for example. A sequence is information indicating a series of modification processes, for example. A plurality of documents belonging to the same sequence ID indicate that they belong to the same series of modification processes. The order information is information indicating the order of modification numerically, and may also be called version information. The content information of the document is information indicating the content of the contract (for example, text, figure, table), for example. The content information of the document is created or uploaded by, for example, a predetermined user who created the draft of the document or a user different from the predetermined user. Different users are other users belonging to the same organization as the predetermined user, for example, another user who can supply document data.

[0023] In the management of documents and the like, there are cases where it is necessary to determine to which sequence a certain document belongs. In particular, in legal documents, since information on the negotiation process may be included in draft documents created during the process of interaction with the other party included in the sequence, there is a strong demand to appropriately manage the file by making it belong to the correct sequence. At this time, it may be possible to determine that the sequence containing other documents determined to be similar in content to the said document is the sequence to which the said document should belong by checking the similarity of the content between the said document and other documents. However, in the case of creating a document using a contract template, for example, it is assumed that there are many documents with very small differences, such as a difference of several characters, from the documents created using the same template. When extracting similar documents from among a plurality of candidates with similar content, the present embodiment that uses the document as a query instead of keywords can be preferably used. Also, legal documents such as such contracts often have large document data, but according to the present embodiment, similar documents can be efficiently extracted.

[0024] In the inverted index database 211, for example, for each user, inverted index information as shown in FIG. 5 is stored. The inverted index information of the present embodiment is information that associates a specific word with the ID of the document containing the specific word for each specific word included in the documents registered in the document database 210. By using this inverted index information, for example, it is possible to specify that the documents containing the word "stock" are the documents with IDs "D0001", "D0003", and "D0010". Also, for example, it is possible to specify that the document containing the word "AAA" is the document with ID "D0001".

[0025] The communication unit 22 of the present embodiment performs various communications with, for example, the terminal device 30. The communication unit 22 acquires, for example, operation information performed on the terminal device 30 by the user. Also, the communication unit 22 transmits various image information for display on the terminal device 30 to the terminal device 30.

[0026] The control unit 23 of this embodiment controls, for example, the document management server device 20. The control unit 23 includes, as a functional configuration realized by executing a program stored in the storage unit 21, a transposed index construction unit 231, a query acquisition unit 232, a similarity calculation unit 233, an extraction unit 234, and an output unit 235.

[0027] (Transposed Index Construction Unit) The transposed index construction unit 231 of this embodiment constructs transposed index information stored in the transposed index database 211. For example, each time a new document is registered in the document database 210 or the content of a document is changed, the transposed index construction unit 231 reads the content information of the document and constructs transposed index information as shown in FIG. 5 from the read content information of the document. Note that, as specific words used in the transposed index information, predetermined words, words specified by the system administrator of the document management server device 20, and the like are used.

[0028] (Query Acquisition Unit) The query acquisition unit 232 of this embodiment acquires a query document that is a document serving as a search query. For example, assume that a user operates a web browser or a predetermined application on the terminal device 30 and designates any one of a plurality of documents stored in the document database 210 and accessible to the user as a search query. In this case, the query acquisition unit 232 acquires the content information of the document designated by the user as the search query from the document database 210. The plurality of documents that the user can access are, for example, in the case of user A shown in FIG. 4, a plurality of documents associated with user A. Hereinafter, the document acquired as the search query by the query acquisition unit 232 is referred to as a "query document".

[0029] Note that the query acquisition unit 232 may acquire a predetermined document input by the user operating a web browser or a predetermined application on the terminal device 30 as the query document.

[0030] In addition, the acquisition of the query document by the query acquisition unit 232 is not limited to being accompanied by an operation by the user, and may be automatically performed by the query acquisition unit 232. For example, when the user causes the terminal device 30 to display a predetermined contract document or uploads a contract document file, the query acquisition unit 232 may automatically set the contract document as the query document and acquire information on the contract document.

[0031] (Similarity calculation unit) The similarity calculation unit 233 of the present embodiment calculates the similarity between the content of the search target document and the content of the query document for each of the plurality of search target documents registered in the document database 210. The similarity calculation unit 233 of the present embodiment calculates the Jaccard similarity between the content of the search target document and the query document. The Jaccard similarity J(x,y) can be calculated by, for example, the following formula f1.

[0032] [Number] In the formula f1, "x" is the set of words included in the search target document. For example, when the search target document contains the content "AAA Co., Ltd.", the set of words x is represented in the form of "x = {stock, company, AAA,...}" by dividing the document into words, for example. "y" is the set of words included in the query document.

[0033] Using the similarity J calculated by the above formula f1, for example, when the similarity J satisfies "J≥t" with respect to a predetermined threshold value t, it can be determined that the content of the search target document is similar to the content of the query document. Also, when "J<t" is satisfied, it can be determined that the search target document is not similar to the query document. The threshold value t is set so as to satisfy "0<t≤1".

[0034] By the way, when the number of search target documents increases, if the above similarity calculation is performed on all of them, the calculation time may become long. On the other hand, among a plurality of search target documents, as shown in FIG. 6, in addition to documents similar to the query document, there are also documents completely different from the query document. Therefore, in order to shorten the similarity calculation time, as shown in FIG. 6, the similarity calculation unit 233 performs a filtering process of narrowing down search target documents estimated to be candidates for the solution of the search result from among a plurality of search target documents, and then performs a similarity calculation only on the narrowed-down search target documents.

[0035] Next, the procedure of the process performed by the similarity calculation unit 233 will be specifically described. First, the similarity calculation unit 233 determines whether the number N of search target documents is equal to or greater than a predetermined number Nth. When the number N of search target documents is less than the predetermined number Nth, the similarity calculation unit 233 extracts search target documents that satisfy the first filtering condition from among the plurality of search target documents as the first search target documents. The first search target documents are search target documents estimated to be candidates for the solution of the search result. The first filtering condition is the condition of satisfying all of the following conditions (a1) to (a3). Note that the "length of the content of a document" is the number of types of words included in a predetermined document. Also, a prefix is a subset from the first word to the predetermined number of words at the beginning of the word set of a predetermined document. (a1) The difference between the length of the content of the query document and the length of the content of the search target document is less than a predetermined length. (a2) At least one of the words included in the prefix of the query document up to the predetermined number and the words included in the prefix of the search target document up to the predetermined number match. (a3) The degree of coincidence between the words included in the prefix of the query document and the words included in the prefix of the search target document is equal to or greater than a predetermined threshold value.

[0036] Here, when the length of the content of a predetermined search target document is denoted as |x| and the length of the content of the query document is denoted as |y|, the similarity calculation unit 233 determines that the condition (a1) is satisfied when the length |x| of the content of the predetermined search target document satisfies the following formula f2 using the above threshold value t.

[0037]

Number

[0038]

Number

[0039]

Number

[0040]

Number

[0041]

Number

[0042] By calculating the similarity after narrowing down the search target documents of the solution candidates in this way, it is possible to shorten the calculation time compared to calculating the similarity for all search target documents. On the other hand, when the number N of search target documents is, for example, 10,000 or more, the number of accesses to the original database increases, and there is a possibility that the search will be slow due to data access. Therefore, when the number of search target documents is particularly large, the similarity calculation unit 233 of the present embodiment further narrows down the search target documents of the solution candidates by using the inverted index information stored in the inverted index database 211.

[0043] Specifically, when the number N of search target documents is equal to or greater than a predetermined number Nth, the similarity calculation unit 233 uses the inverted index information to extract, from among the plurality of search target documents, the search target documents that satisfy the second filtering condition as the second search target documents. The second search target documents are the search target documents that are estimated to be candidates for the solutions of the search results. The second filtering condition is the condition that specific words included in the query document are included in the search target documents. For example, when a character string from the first character to the nth character at the head of a predetermined document is used as a prefix, the similarity calculation unit 233 extracts some words included in the prefix of the query document as search words, and then extracts, based on the inverted index information, the search target documents that include the extracted search words. Specifically, as shown in FIG. 7, when words such as "{stock, company, XXX,...}" are included in the prefix of the query document, the similarity calculation unit 233 extracts "stock" and "XXX" as search words. Then, the similarity calculation unit 233 uses the inverted index information shown in FIG. 5 to extract, as solution candidates, the documents with document IDs "D0001", "D0003", and "D0010" from the word "stock". Subsequently, the similarity calculation unit 233 executes a process of calculating the similarity for the extracted single or multiple solution candidate search target documents.

[0044] (Extraction unit) The extraction unit 234 of the present embodiment extracts, from among the plurality of search target documents, the search target documents with a high similarity to the query document as the search result documents.

[0045] For example, when the number N of search target documents is less than the predetermined number Nth, the similarity calculation unit 233 extracts a plurality of first search target documents and calculates the similarity of each of the plurality of first search target documents. In this case, the extraction unit 234 extracts the first search target documents that satisfy "J≧t" in terms of the similarity J as the search result documents with a high similarity to the query document.

[0046] On the other hand, when the number N of documents to be searched is equal to or greater than a predetermined number Nth, the similarity calculation unit 233 extracts a plurality of second documents to be searched and calculates the similarity of each of the plurality of second documents to be searched. In this case, the extraction unit 234 extracts, as search result documents with a high similarity to the query document, the second documents to be searched for which the similarity J satisfies "J≧t".

[0047] (Output unit) The output unit 235 of the present embodiment transmits information related to the single or multiple search result documents extracted by the extraction unit 234 to the terminal device 30, thereby causing the terminal device 30 to display the information related to the search result documents through a web browser or a predetermined application. The information related to the search result documents may be the title or content information of the search result documents themselves, or may be the title or content information of other documents related to the search result documents. The other documents have, for example, the same sequence ID as the search result documents and are the documents with the latest sequence information. Further, the information related to the search result documents may be the titles and content information of the plurality of documents corresponding to each of the plurality of sequence information, each having the same sequence ID as the search result documents, and the editing history of those documents.

[0048] (Operation example of the document processing system) Next, an operation example of the document processing system 10 of the present embodiment will be described.

[0049] As shown in FIG. 8, in the document processing system 10 of the present embodiment, first, it is assumed that the user operates the terminal device 30 to perform an operation of specifying a query document (step S20). Then, the query acquisition unit 232 of the document management server device 20 acquires the query document specified by the user (step S30). As a result, the similarity calculation unit 233 of the document management server device 20 calculates the similarity of each of the plurality of documents to be searched registered in the document database 210 with respect to the query document. Then, the extraction unit 234 of the document management server device 20 extracts, as search result documents, the documents to be searched with a high similarity to the query document from among the plurality of documents to be searched.

[0050] Specifically, the similarity calculation unit 233 determines whether the number N of search target documents is equal to or greater than a predetermined number Nth (step S31). If the number N of search target documents is less than the predetermined number Nth (step S31: NO), a plurality of search target documents that satisfy the first filtering condition are extracted from the plurality of search target documents as the first search target documents (step S32), and the similarity of each of the extracted plurality of first search target documents is calculated (step S33). Then, the extraction unit 234 extracts, from the plurality of first search target documents, a first search target document whose similarity J satisfies "J ≧ t" as a search result document with a high similarity to the query document (step S34).

[0051] On the other hand, when the number N of search target documents is equal to or greater than the predetermined number Nth (step S31: YES), the similarity calculation unit 233 extracts a plurality of search target documents that satisfy the second filtering condition from the plurality of search target documents as the second search target documents (step S35), and calculates the similarity of each of the extracted plurality of second search target documents (step S36). Then, the extraction unit 234 extracts, from the plurality of second search target documents, a second search target document whose similarity J satisfies "J ≧ t" as a search result document with a high similarity to the query document (step S37).

[0052] Subsequently, the output unit 235 of the document management server device 20 transmits information related to the search result document extracted by the process of step S34 or step S37 to the terminal device 30 (step S38). Thereby, the terminal device 30 displays the information related to the search result document (step S21).

[0053] Thereafter, the user performs a predetermined operation. For example, the user can associate the query document or a document related to the query document with the sequence to which the search result document belongs according to a proposal from the system (a type of information related to the search result document). In that case, the query document may be recorded and treated as the document (latest version) with the latest sequence information of the documents in the sequence based on information such as the upload date and update date information.

[0054] (Hardware Configuration of the Document Processing System) Next, with reference to FIG. 9, an example of the hardware configuration when the document management server device 20 and the terminal device 30 are realized by the computer 1000 will be described. FIG. 9 is a diagram showing an example of the hardware configuration of the computer 1000.

[0055] As shown in FIG. 9, the computer 1000 includes, for example, a processor 1001, a memory 1002, a storage device 1003, an input I / F unit 1004, a data I / F unit 1005, a communication I / F unit 1006, and a display device 1007. Note that the computer 1000 may include a plurality of the processor 1001, the memory 1002, the storage device 1003, the input I / F unit 1004, the data I / F unit 1005, the communication I / F unit 1006, and the display device 1007, respectively.

[0056] The computer 1000 may be, for example, a server computer, a personal computer (e.g., desktop, laptop, tablet, etc.), a media computer platform (e.g., cable, satellite set-top box, digital video recorder, etc.), a handheld computer device (e.g., PDA, email client, etc.), or another type of computer, or a communication platform.

[0057] The processor 1001 is, for example, a control unit that controls various processes in the computer 1000 by executing a program stored in the memory 1002.

[0058] The memory 1002 is a storage medium such as a RAM (Random Access Memory), for example. The memory 1002 temporarily stores the program code of the program executed by the processor 1001 and the data required during the execution of the program.

[0059] The storage device 1003 is a non-volatile storage medium such as a hard disk drive (HDD) or a flash memory. The storage device 1003 stores an operating system and various programs for implementing the above-described respective configurations.

[0060] The input I / F unit 1004 is a device for receiving an input from a user. The input I / F unit 1004 is, for example, a keyboard, a mouse, a touch panel, various sensors, a wearable device, or the like. The input I / F unit 1004 may be connected to the computer 1000 via an interface such as USB (Universal Serial Bus).

[0061] The data I / F unit 1005 is a device for inputting data from outside the computer 1000. The data I / F unit 1005 is, for example, a drive device for reading data stored in various storage media. The data I / F unit 1005 may be provided outside the computer 1000. When the data I / F unit 1005 is provided outside the computer 1000, the data I / F unit 1005 is connected to the computer 1000 via an interface such as USB.

[0062] The communication I / F unit 1006 is a device for performing data communication via a network such as the Internet, either wired or wirelessly, with a device outside the computer 1000. The communication I / F unit 1006 may be provided outside the computer 1000. When the communication I / F unit 1006 is provided outside the computer 1000, the communication I / F unit 1006 is connected to the computer 1000 via an interface such as USB.

[0063] The display device 1007 is a device for displaying various information. The display device 1007 is, for example, a liquid crystal display, an organic EL (Electro-Luminescence) display, a display of a wearable device, or the like. The display device 1007 may be provided outside the computer 1000. When the display device 1007 is provided outside the computer 1000, the display device 1007 is connected to the computer 1000 via, for example, a display cable or the like. Further, when a touch panel is adopted as the input I / F unit 1004, the display device 1007 may be configured integrally with the input I / F unit 1004.

[0064] (Operations and Effects of the Document Processing System of the Present Embodiment) As described above, the document processing system 10 of the present embodiment includes a query acquisition unit 232, a similarity calculation unit 233, an extraction unit 234, and an output unit 235. The query acquisition unit 232 acquires a query document that is a document serving as a search query. The similarity calculation unit 233 calculates the similarity of the content of each of a plurality of search target documents, which are documents to be searched based on the query document, to the content of the query document. The extraction unit 234 extracts, from among the plurality of search target documents, a search target document having a high similarity to the query document based on the similarity of each of the plurality of search target documents. The output unit 235 outputs information related to the search target document having a high similarity to the query document, which is extracted by the extraction unit 234.

[0065] According to this configuration, if the user designates a query document, a search target document having a high similarity to the content of the query document is extracted, and information related to the search target document is output. Therefore, it is possible to search for documents having similar contents.

[0066] When calculating the similarity, if the number N of search target documents is less than a predetermined number Nth, the similarity calculation unit 233 extracts, as first search target documents, the search target documents that satisfy the first filtering condition from among the plurality of search target documents, and calculates the similarity of the extracted first search target documents. The first filtering condition is the condition of satisfying all of the above conditions (a1) to (a3).

[0067] According to this configuration, when the number N of search target documents is less than the predetermined number Nth, after the search target documents that satisfy the first filtering condition are narrowed down as the first search target documents from among the plurality of search target documents, the similarity of the narrowed-down first search target documents is calculated. Therefore, compared with the case of calculating the similarity of all search target documents, it is possible to shorten the calculation time of the similarity.

[0068] When calculating the similarity, if the number N of search target documents is greater than or equal to the predetermined number Nth, the similarity calculation unit 233 reads the inverted index information from the storage unit 21, and based on the inverted index information, extracts, as second search target documents, the search target documents that include the words of the query document from among the plurality of search target documents, and calculates the similarity of the second search target documents.

[0069] According to this configuration, when the number N of search target documents is greater than or equal to the predetermined number Nth, after the search target documents that satisfy the second filtering condition are narrowed down as the second search target documents from among the plurality of search target documents, the similarity of the narrowed-down second search target documents is calculated. Therefore, compared with the case of calculating the similarity of all search target documents, it is possible to shorten the calculation time of the similarity.

[0070] (Other embodiments) The present disclosure is not limited to the above specific examples.

[0071] For example, the documents targeted by the document processing system 10 are not limited to contract documents, and may be official documents, minutes of meetings, rules, notices, regulations, and regulations.

[0072] The first filtering condition and the second filtering condition can be changed as appropriate. For example, the first filtering condition may be a condition that satisfies at least one of the conditions (a1) to (a3) described above.

[0073] In the process of step S34 shown in FIG. 8, the output unit 235 may extract a predetermined number of first search target documents with the highest similarity J from among a plurality of first search target documents as search result documents with a high similarity to the query document. Further, in the process of step S37 shown in FIG. 8, the output unit 235 may execute the same process for the second search target document.

[0074] The functional configurations respectively possessed by the document management server device 20 and the terminal device 30 may be provided only in the document management server device 20, only in the terminal device 30, or in each of the document management server device 20 and the terminal device 30. For example, the inverted index construction unit 231, the query acquisition unit 232, the similarity calculation unit 233, the extraction unit 234, and the output unit 235 possessed by the document management server device 20 may be provided in the terminal device 30. Further, the document processing system 10 may include a plurality of the document management server device 20 and the terminal device 30 respectively.

[0075] Those in which those skilled in the art have appropriately made design changes to the above specific examples are also included in the scope of the present disclosure as long as they have the features of the present disclosure. Each element included in each of the above-described specific examples, and its arrangement, conditions, shape, etc. are not limited to those illustrated and can be changed as appropriate. Each element included in each of the above-described specific examples can be combined as appropriate without causing a technical contradiction.

Explanation of Reference Numerals

[0076] 10: Document processing system, 232: Query acquisition unit, 233: Similarity calculation unit, 234: Extraction unit, 235: Output unit, 1001: Processor.

Claims

1. A processor obtains a query document that is a document serving as a search query, calculates, for each of a plurality of search target documents that are documents to be searched based on the query document, a similarity between the content of the search target document and the content of the query document, extracts, from among the plurality of search target documents, a search target document having a high similarity to the query document based on the similarity of each of the plurality of search target documents, and outputs information related to the extracted search target document having a high similarity to the query document. A document processing method.

2. When calculating the similarity, if the number of the search target documents is less than a predetermined number, the processor extracts, from among the plurality of search target documents, a search target document that satisfies a predetermined filtering condition as a first search target document, and calculates the similarity of the first search target document. The document processing method according to claim 1.

3. The processor uses, as the predetermined filtering condition, a condition that a difference between the length of the content of the query document and the length of the content of the search target document is less than a predetermined length. The document processing method according to claim 2.

4. When a subset from the first word to the nth word of the word set of a predetermined document is used as a prefix, the processor uses, as the predetermined filtering condition, a condition that at least one word included in the prefix of the query document up to the nth word matches a word included in the prefix of the search target document up to the nth word. The document processing method according to claim 2.

5. When a subset from the first word to the nth word of the word set of a predetermined document is used as a prefix, the processor uses, as the predetermined filtering condition, a condition that a degree of coincidence between the words included in the prefix of the query document and the words included in the prefix of the search target document is equal to or more than a predetermined threshold value. The document processing method according to claim 2.

6. When calculating the similarity, if the number of the search target documents is equal to or more than a predetermined number, the processor reads, from a storage unit, inverted index information capable of specifying a search target document including the word from the words included in each of the plurality of search target documents, extracts, from among the plurality of search target documents, a second search target document including the word of the query document based on the inverted index information, and calculates the similarity of the second search target document. ​ ​ ​ ​ ​ ​ The document processing method according to claim 1.

7. The processor calculates the Jaccard similarity as the similarity The document processing method according to claim 1.

8. The processor obtains the plurality of search target documents from a database The document processing method according to claim 1.

9. The query document and the search target documents are contract documents The document processing method according to claim 1.

10. A processor is caused to perform a process of obtaining a query document which is a document serving as a search query, and for each of the plurality of search target documents which are documents searched based on the query document, perform a process of calculating a similarity between the content of the search target document and the content of the query document with respect to the content of the query document, and based on the similarity of each of the plurality of search target documents, perform a process of extracting a search target document having a high similarity to the query document from among the plurality of search target documents, and perform a process of outputting information related to the search target document having a high similarity to the query document, which is extracted program.

11. A query acquisition unit that acquires a query document which is a document serving as a search query, a similarity calculation unit that calculates a similarity between the content of the search target document and the content of the query document with respect to the content of the query document for each of the plurality of search target documents which are documents searched based on the query document, an extraction unit that extracts a search target document having a high similarity to the query document from among the plurality of search target documents based on the similarity of each of the plurality of search target documents calculated by the similarity calculation unit, and an output unit that outputs information related to the search target document having a high similarity to the query document, which is extracted by the extraction unit document processing system.

Citation Information

Patent Citations

  • Document processing program, information processing device and document processing method

    JP2022114897A