Document search system
The method improves document retrieval by dividing documents into sentence blocks, calculating relevance and similarity scores, addressing inefficiencies in identifying similar documents, especially for intellectual property, enhancing accuracy and reducing processing time.
Patent Information
- Application Number
- JP2025109022
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-11-30
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-17
- Estimated Expiration
- 2039-11-19
AI Technical Summary
Existing document search methods struggle to accurately identify similar documents at a granular level, such as sentence blocks, leading to inefficiencies in document retrieval, especially in the context of intellectual property creation, where precise reference and citation are crucial.
A document retrieval method that divides documents into sentence blocks, performs full-text searches using these blocks as criteria, calculates relevance and similarity scores, and selects similar blocks based on these scores to enhance accuracy and efficiency.
Enables high-accuracy retrieval of similar documents at a sentence block level, reducing processing time and user input complexity, particularly beneficial for intellectual property searches.
Smart Images

Figure 2025134970000001_ABST
Abstract
Description
[Technical Field]
[0001] One aspect of the present invention is a document search method, a document search system, a program, and a non-transitory computer The present invention relates to a computer-readable storage medium.
[0002] Note that one embodiment of the present invention is not limited to the above technical field. Examples of the semiconductor device include a semiconductor device, a display device, a light-emitting device, a power storage device, a memory device, an electronic device, a lighting device, Input devices (e.g., touch sensors), input / output devices (e.g., touch panels), etc. These driving methods or manufacturing methods can be cited as examples. [Background technology]
[0003] Document search technology that can efficiently search for a target document from a large number of documents is being actively developed. For example, Patent Document 1 discloses a similar document search method.
[0004] Similar documents may be similar to the target document overall, but may have extremely similar parts. In some cases, the similarity is high in some areas, while in other areas the similarity is extremely low.
[0005] In Patent Document 1, similar documents are compared to a target document to determine whether they are similar overall or only partially. The level of detail is calculated as an index to determine how similar the items are. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Japanese Patent Application Laid-Open No. 2004-295712 Summary of the Invention [Problem to be solved by the invention]
[0007] In patent application work, when creating a new specification (specification of a later application), The description of the specification (specification of the prior application) may be referred to or cited. If a translation of the specification of the earlier application has already been prepared, when preparing a translation of the specification of the later application, The translation of the specification of the earlier application can be used as a reference or cited, and the translation of the specification of the later application can be used as a reference or cited. This can reduce the time required.
[0008] Depending on the search method for similar documents, documents that have a high degree of similarity to the target document may be found. Even if they are not actually similar, they have a certain degree of similarity overall, so the classification of the entire document is Some documents may be highly similar, while the rest may be very low similar. Even if the document has extremely high similarity (e.g., contains a perfect match), For example, if you refer to a translation, or The latter document is preferred over the former for citation.
[0009] Also, by searching sentence by sentence, you can find exact matches, but The flow of the text may be interrupted, and translations may not be consistent across specifications. Therefore, it is desirable to be able to grasp similar parts in units of sentences containing multiple sentences, such as by chapter. .
[0010] In addition, there is not necessarily only one specification to refer to when creating a new specification. Not only which specification was used as a reference to create the new specification, but also which part of which specification It is desirable to be able to easily understand which parts of the new specification have been prepared by referring to the previous part. This is true not only for specifications, but for all documents. When creating a new document, keep a detailed record of which documents and which parts of them you referred to. It is a time-consuming and complicated task.
[0011] One aspect of the present invention provides a document retrieval method that can retrieve similar documents for each block of a document. Another object of the present invention is to provide a method for classifying blocks of a document. One of the objectives is to provide a document search system that can search for similar documents. One aspect of the present invention is to enable a search for similar documents for each block of a document using a simple input method. One of the objectives of the present invention is to provide a document search method that can
[0012] An object of one aspect of the present invention is to provide a document retrieval method that can retrieve documents with high accuracy. Another aspect of the present invention is a document retrieval system that can retrieve documents with high accuracy. Another object of the present invention is to provide a method for inputting a simple and accurate One of the goals is to achieve high-performance document search, especially for documents related to intellectual property.
[0013] Note that the description of these problems does not preclude the existence of other problems. It is not necessary to solve all of these problems. From the description of the section, it is possible to extract other issues. [Means for solving the problem]
[0014] One aspect of the present invention is to provide a method for dividing a plurality of documents to be searched into a plurality of sentences. A document retrieval method for retrieving a specific sentence block from among blocks, A first search text block is prepared, which is a part of the plurality of text blocks. A part of the text is treated as the first target, and a full-text search is performed using the first search text block as the search criteria. By doing so, for each sentence block included in the first target, and calculating a first relevance degree for the first target, and selecting a second target from the first target based on the degree of the first relevance degree. The target is determined, and for each sentence in the first search sentence block, the sentences in the second target are A first similarity is calculated for each of the first and second search sentence blocks. This is a document retrieval method that searches for at least one block of text that is similar to
[0015] It is preferable to divide the search document into multiple search text blocks. In this case, the first search text block is preferably one of the plurality of search text blocks. Desirable.
[0016] Furthermore, a second search text block, which is another part of the search document, is prepared, and a plurality of sentences are The second search text block is used as a search criterion, with at least part of the block as the third target. By performing a full-text search using the third target, the second Calculate a second relevance for the search sentence block, and based on the second relevance, The fourth object is determined from the third object, and for each sentence included in the second search sentence block, , a second similarity with each of the sentences included in the fourth target is calculated, and the second similarity is used to It is preferred to search for at least one block of text similar to the second search block of text. In this case, the first object and the third object may be the same or different from each other. That's fine.
[0017] Using a value of the first similarity equal to or greater than a threshold, a sentence block similar to the first search sentence block is selected. It is preferable to search for at least one lock.
[0018] In one aspect of the present invention, for each of a plurality of search sentence blocks, a plurality of search target documents are From the multiple sentence blocks created by dividing each sentence, similar sentence blocks are selected. A document retrieval method for retrieving a document, which divides a document to be searched into a plurality of text blocks to be searched. Create a lock and for each of the multiple search text blocks, At least a part of the sentences is used as the first target, and a full-text search is performed using the search sentence block as a search condition. By searching, the search text block for each text block included in the first target is calculating a degree of relevance to the first object; and selecting a second object from the first object based on the degree of relevance. determining a target of the second target, and for each sentence included in the search sentence block, determining a target of the second target of the second target. a step of calculating a similarity between each of the sentences included in the search block and the search text block using the similarity; and searching for at least one block of text similar to the block. is.
[0019] One aspect of the present invention is to provide a method for dividing a plurality of documents to be searched into a plurality of sentences. This is a document search method that searches for a specific sentence block from among the blocks, and A first search sentence block is prepared, which is a part of the plurality of sentence blocks. The first search block is used to search the entire text. By performing a sentence search, each sentence included in the first target is searched for in the first sentence block. A first relevance score is calculated for each sentence included in the first search sentence block. In each case, the second target is determined from the sentences included in the first target based on the degree of relevance of the first target. For each sentence in the first search sentence block, A first similarity is calculated with the first search sentence block, and the first similarity is used to find a sentence similar to the first search sentence block. This is a document retrieval method that searches for at least one block of text that matches the given text.
[0020] It is preferable to divide the search document into multiple search text blocks. In this case, the first search text block is preferably one of the plurality of search text blocks. Desirable.
[0021] Furthermore, a second search text block, which is another part of the search document, is prepared, and a plurality of sentences are At least a part of the block is treated as a third target, and the second search sentence block is included in the By performing a full-text search using each sentence in the third target as a search condition, A second relevance score is calculated for each sentence included in the second search sentence block, and the second search sentence block is calculated. For each sentence in the sentence block, the sentences included in the third target are selected based on the second relevance. A fourth target is determined from the sentences included in the second search sentence block, and the fourth target is determined from the sentences included in the second search sentence block. A second similarity is calculated for each sentence included in the target, and the second similarity is used to It is preferable to search for at least one text block similar to the search text block. In this case, the first object and the third object may be the same or different from each other. stomach.
[0022] Using a value of the first similarity equal to or greater than a threshold, a sentence block similar to the first search sentence block is selected. It is preferable to search for at least one lock.
[0023] In one aspect of the present invention, for each of a plurality of search sentence blocks, a plurality of search target documents are From the multiple sentence blocks created by dividing each sentence, similar sentence blocks are selected. A document retrieval method for retrieving a document, which divides a document to be searched into a plurality of text blocks to be searched. Create a lock and for each of the multiple search text blocks, At least a part of the sentences in the search block is used as the first target. By using this to perform a full-text search, the search sentence blocks for each sentence included in the first target are obtained. a step of calculating the relevance of each sentence included in the search sentence block; For each sentence, determine the second target from among the sentences contained in the first target based on the degree of relevance. and for each sentence included in the search sentence block, searching for each sentence included in the second target. A step of calculating the similarity between each of the sentence blocks, and using the similarity, a step of calculating the similarity between each of the sentence blocks, and a step of calculating the similarity between each of the sentence blocks, and retrieving at least one block of text.
[0024] One aspect of the present invention is a document search system having a function of performing any of the above document search methods. is.
[0025] One aspect of the present invention is to provide a method for dividing a plurality of documents to be searched into a plurality of sentences. A document retrieval system for retrieving a specific sentence block from among blocks, comprising a processing unit The processing unit divides the document to be searched into a plurality of search sentence blocks. The first function is to prepare a text block for search, and the second function is to select at least one of the multiple text blocks. A full-text search is performed using at least a portion of the text as the first target and the first search text block as the search criteria. By performing this, the first search sentence block for each sentence block included in the first target is obtained. A function for calculating a first relevance to the target based on the degree of the first relevance. The function of determining the second target from among the above, and for each sentence contained in the first search sentence block, A function for calculating a first similarity with each sentence included in the second target, and a function for calculating a first similarity with each sentence included in the second target using the first similarity. a function of searching for at least one block of text similar to the first block of text for search; and a document retrieval system having:
[0026] One aspect of the present invention is a method for causing a processor to execute any of the document search methods described above. One aspect of the present invention is a non-transitory computer program that stores the program. It is a data-readable storage medium.
[0027] The program can be stored in a computer by various types of temporary computer-readable storage media. The temporary computer-readable storage medium may include an electrical signal, an optical signal, and electromagnetic waves. Transient computer-readable storage media include wired storage media such as electrical wires and optical fibers. The program can be supplied to the computer via a communication channel or a wireless communication channel.
[0028] One aspect of the present invention is to provide a method for dividing a plurality of documents to be searched into a plurality of sentences. A program that searches for a specific text block from among the blocks, and divides the search document The first search sentence block is one of the multiple search sentence blocks created by dividing preparing a block; and selecting at least a portion of the plurality of sentence blocks as a first target. Then, by performing a full-text search using the first search sentence block as a search condition, the first target is found. Calculate the first relevance of each included text block to the first search text block. and determining a second target from among the first targets based on the degree of first relevance. and for each sentence included in the first search sentence block, determining whether the second target sentence is included in the first search sentence block. calculating a first similarity between each of the sentences, and performing a first search using the first similarity; searching for at least one sentence block similar to the searched sentence block; In one aspect of the present invention, the program is stored in a It is a non-transitory computer-readable storage medium.
[0029] Non-transitory computer-readable storage media include various types of tangible storage media. The non-transitory computer-readable storage medium may be, for example, a RAM (Random Access Memory). volatile memory such as ROM (Read Only Access Memory), Non-volatile memory such as hard disk drives (HDDs) is also included. (Hard Disc Drive: HDD) and Solid State Drive (Soli Recording media drives such as SSDs, magneto-optical disks, Examples include DVD-ROMs and CD-Rs. [Effects of the Invention]
[0030] According to one aspect of the present invention, a document retrieval method capable of retrieving similar documents for each block of a document is provided. According to one aspect of the present invention, it is possible to search for similar documents for each block of a document. According to one aspect of the present invention, a document search system can be provided that allows a user to search for documents by a simple input method. For each lock, a document search method can be provided that can search for similar documents.
[0031] According to one aspect of the present invention, a document retrieval method capable of retrieving documents with high accuracy can be provided. According to one aspect of the present invention, a document retrieval system capable of retrieving documents with high accuracy can be provided. According to one aspect, a highly accurate document search, particularly a search for documents related to intellectual property, can be performed using a simple input method. This can be achieved.
[0032] The description of these effects does not preclude the existence of other effects. However, it is not necessary to have all of these effects. , it is possible to extract effects other than these. [Brief explanation of the drawings]
[0033] [Figure 1] FIG. 1 is a flow diagram showing an example of a document retrieval method. [Figure 2] FIG. 2 is a diagram showing an example of a process before a search is performed. [Figure 3] 3A, 3B, and 3C are diagrams showing an example of a document search method. [Figure 4] 4A, 4B, and 4C are diagrams showing an example of a document search method. [Figure 5] 5A and 5B are diagrams showing an example of a document search method. [Figure 6] 6A, 6B, and 6C are diagrams showing an example of a document search method. [Figure 7] 7A, 7B, and 7C are diagrams showing an example of a document search method. [Figure 8] 8A, 8B, and 8C are diagrams showing an example of a document search method. [Figure 9] 9A and 9B are diagrams showing an example of a document search method. [Figure 10] FIG. 10 is a flow diagram showing an example of a document search method. [Figure 11] FIG. 11 is a flow diagram showing an example of a document search method. [Figure 12] FIG. 12 is a diagram showing an example of a document search method. [Figure 13] FIG. 13 is a block diagram showing an example of a document search system. [Figure 14] FIG. 14 is a block diagram showing an example of a document search system. DETAILED DESCRIPTION OF THE INVENTION
[0034] The embodiments will be described in detail with reference to the drawings. However, the present invention is not limited to the following description. The present invention is not limited to the above embodiments, and various changes and modifications may be made in the form and details thereof without departing from the spirit and scope of the present invention. It will be readily understood by those skilled in the art that the present invention can be achieved by the following embodiments. It should not be construed as being limited to the contents described.
[0035] In the configuration of the invention described below, the same parts or parts having similar functions are designated by the same reference numerals. The same reference numerals are used in common among different drawings, and the repeated explanations thereof will be omitted. When referring to a function, the hatch pattern may be the same and no particular symbol may be added.
[0036] In addition, the position, size, range, etc. of each component shown in the drawings are not necessarily the same as in reality for ease of understanding. Therefore, the disclosed invention may not necessarily represent the position, size, range, etc. Furthermore, the present invention is not limited to the position, size, range, etc. disclosed in the drawings.
[0037] (Embodiment 1) In this embodiment, a document search method according to one embodiment of the present invention will be described with reference to FIGS. Note that the data diagram is an example and is not limiting.
[0038] One aspect of the present invention is to provide a method for dividing a plurality of documents to be searched into a plurality of sentences. This is a document retrieval method that searches for a specific sentence block within a block.
[0039] First, a first search sentence block, which is a part of a search document, is prepared.
[0040] For example, the first search sentence block can be created by extracting a part of the search document. Alternatively, the first search sentence block may be a plurality of search sentences created by dividing the search document. It may also be one of the search text blocks.
[0041] In a document retrieval method according to one aspect of the present invention, a plurality of sentence blocks are extracted in advance from a plurality of documents to be retrieved. In addition, when searching, a search block is created from the search document. This allows you to search for text blocks that are similar to the search text block. Therefore, when the entire document is used as a search condition or when the search target is the entire document, By comparing these, it becomes easier to understand the correspondence between similar parts.
[0042] Next, at least a part of the plurality of sentence blocks is used as a first target, and a first search sentence is By performing a full-text search using the block as a search condition, the sentence blocks included in the first target are found. A first relevance level for each of the first search sentence blocks is calculated.
[0043] The more documents to be searched, the more sentence blocks there will be. For each text block, you can narrow down the text block to be searched (first target). This reduces the amount of processing and increases search speed.
[0044] Next, a second target is determined from among the first targets based on the level of the first relevance.
[0045] In full-text search, the order of sentences and words is not taken into account, so the calculated relevance is different from the similarity. On the other hand, sentence blocks that share words with the search sentence block have a relevance value of The similarity is calculated based on the similarity of the sentence blocks. This allows you to narrow down the target with high precision.
[0046] Next, for each sentence in the first search sentence block, A first similarity is calculated between the
[0047] The process of calculating the similarity tends to take longer than the full-text search. In this method, the second object is determined from the first object, and the similarity is calculated after narrowing down the objects. This reduces the time required for document search.
[0048] Similarity can be calculated based on the degree of similarity between the texts of the sentences. The order of words in a sentence is taken into consideration when calculating the similarity. A sentence that has common words with a sentence in a chapter block but has a different word order is considered to have a similarity score. becomes lower.
[0049] Then, using the first similarity, a small number of sentence blocks similar to the first search sentence block are selected. Find at least one.
[0050] As described above, by using the document search method according to one embodiment of the present invention, it is possible to search for a specific part of a document to be searched. It is easy to find similar references in other documents.
[0051] In addition, in the document search method according to one aspect of the present invention, it is sufficient to input a search document, and the keywords to be used for the search are Since there is no need to select keywords, the burden on the user is small and the search results do not differ depending on the user's skill. It has the advantage of being less susceptible to damage.
[0052] In addition, after narrowing down the text blocks to be searched to the first target, then the second target, Since similarity is calculated, the time required for document search can be reduced.
[0053] In addition, the full-text search uses each sentence contained in the first search sentence block as a search condition. In this case, a first search sentence block for each sentence included in the first target is generated. A first relevance score is calculated for each sentence in the first search sentence block. For each sentence in the table, select a sentence from the first target based on the degree of first relevance. Determine the second target.
[0054] A sentence block contains multiple sentences. Among the sentences in the sentence block, the first search The majority of sentences in a sentence block are not necessarily similar to those in the other sentences. In order to search for high-quality text blocks with high accuracy, it is necessary to It may take a long time to calculate the similarity. The second objective is to reduce the number of sentence blocks in order to reduce the time required for extraction. This means that there is a risk of missing sentence blocks that contain highly similar sentences.
[0055] Therefore, we narrow down the second target from the first target by sentence, not by sentence block. Specifically, for each sentence included in the first search sentence block, a highly relevant It is preferable to search for sentences and narrow down the targets for which similarity is to be calculated on a sentence-by-sentence basis. By narrowing down the target, sentences with high similarity ( and sentence blocks) and shorten the time required to calculate similarity. This can be achieved.
[0056] <Document search method example 1> FIG. 1 shows a flowchart of a document retrieval method. As shown in FIG. 1, a document retrieval method according to an embodiment of the present invention is The book search method includes six steps, Step A1 to Step A6.
[0057] Unless otherwise specified, a structure that has multiple elements (documents, text blocks, sentences, etc.) Even when explaining the composition of a product, if you explain matters common to each element, you should use variables and For example, the search target document TD1 and the search target document TD2 are When explaining matters common to the search target document TDn, etc., we will refer to the search target document TD. There are cases where this happens.
[0058] [Preprocessing] First, the process before the search will be described with reference to FIG.
[0059] In the preprocessing, multiple search target documents TD are divided into multiple text blocks TB.
[0060] In the document search method of this embodiment, a plurality of documents prepared in advance are divided into blocks. Then, when searching, the input search document is also divided into blocks. It is possible to search for sentence blocks similar to each block.
[0061] FIG. 2 shows an example in which n (n is an integer of 2 or more) documents TD to be searched are prepared.
[0062] There are no particular limitations on the documents TD to be searched, and various documents can be used.
[0063] The search target document TD may be, for example, a document related to intellectual property. Specifically, the documents include the specification, claims, and abstract used in the patent application. Furthermore, documents related to intellectual property include patent documents (unpublished patent publications, patent publications, etc.). These include publications such as patent gazettes, utility model gazettes, design gazettes, and papers. Publications issued in countries around the world, not just those published in Japan, can be used as documents related to intellectual property. It is possible.
[0064] In addition, the search target documents TD can be books, papers, reports, columns, or other Various works including texts may be used. In addition, medical records may be used as search target documents. It's fine.
[0065] There are also no particular restrictions on the language of the document, and it can be in Japanese, English, Chinese, Korean, etc. Which document can be used?
[0066] The search target document TD1 shown in Figure 2 is divided into x (x is an integer equal to or greater than 2) text blocks (text blocks The text is divided into blocks TB1(1) to TB1(x).
[0067] The search target document TD2 is composed of y (y is an integer of 2 or more) text blocks (text blocks TB2(1) is divided into sentence blocks TB2(y)).
[0068] In addition, the search target document TDn is made up of z (z is an integer of 2 or more) text blocks (text blocks It is divided into sentence blocks TBn(1) to TBn(z).
[0069] For example, if the document to be searched consists of multiple chapters, dividing it into chapters will reduce the number of You may create a number of sentence blocks.
[0070] Specifically, in the case of a patent specification, "Background, Problems, Means, and Effects," "Embodiment 1," It can be divided into "Embodiment 2" and so on.
[0071] In addition, in the case of a thesis, it should be divided into sections such as "Introduction," "Research Method," "Results," "Discussion," and "Conclusion." It is possible.
[0072] It is also possible to create multiple sentence blocks using all sentences in the document to be searched. Multiple text blocks may be created using only the necessary parts of the target document.
[0073] For example, if the document to be searched is a patent specification, multiple sentence blocks can be searched without using "symbol explanations." You may also create a block.
[0074] The preprocessing is performed at least once before the document search (before step A1 is performed). The preprocessing may be performed multiple times depending on the purpose. For example, preprocessing may be performed periodically to Adding, updating, or deleting documents can improve search accuracy and usability. .
[0075] Furthermore, multiple text blocks are used to create an index file for full-text search. It is preferable to create a rule that allows full-text searches to be performed quickly. The structure of the index file is not particularly limited, and may include, for example, a character string, a document name, a text block name, etc. , frequency of occurrence, etc.
[0076] For example, the index file contains the search target document TD (or text block TB) The information may include whether or not there is a translation in each language. Specifying conditions such as "English translation exists" or "Chinese translation exists" can be done.
[0077] Next, the six steps shown in FIG. 1 will be described in detail with reference to FIGS.
[0078] [Step A1: Creating multiple search text blocks STB] First, the search document STD is divided into multiple search text blocks STB. (Figure 3A).
[0079] As shown in FIG. 3A, the search document STD is composed of w search sentence blocks (w is an integer of 2 or more). block (search text block STB(1) to search text block STB(w)) can be.
[0080] In the document retrieval method of this embodiment, the input search document STD is divided into a plurality of search text blocks. In order to separate the text block STBs into search text block STBs, similar documents (text block STBs) are You can search for the TB.
[0081] There are no particular limitations on the search document STD, and various documents can be used.
[0082] An example of a search document STD is a document related to intellectual property before translation. This allows you to search for similar documents that have already been translated from the target document TD, and Able to refer to or quote passages.
[0083] In addition, as search document STD, books, papers, reports, columns, or various documents containing sentences are available. This allows you to search for similar documents from the search target documents TD. It is possible to search and check whether the search document STD is free from any suspicion of plagiarism or plagiarism. can be done.
[0084] In addition, medical records can be used as search documents STD. By using the collected medical records, medical records of similar cases can be searched for and used as a reference for medical treatment. It is also possible to consider how the patient will progress in the future.
[0085] [Step A2: Selection of search text block STB(i)] Next, from among the w search sentence blocks STB, the search sentence block STB to be searched is selected. Select (i) (i is an integer between 1 and w).
[0086] If you want to search only one search text block STB, go to step A1. By extracting necessary parts from the search document STD, the search text block ST You can also create B.
[0087] Also, when searching for multiple search text blocks STB, you need to search them one by one. You can search multiple documents in parallel (see Example 3 of document search method). (See Example 4 of Search Method) and search may be performed by combining sequential processing and parallel processing.
[0088] In the document search method of this embodiment, for each search text block STB, similar text blocks are searched. It is possible to search the TB to find search targets that are similar to specific parts of the search document STD. It is possible to accurately and easily grasp the written location of the target document TD.
[0089] [Step A3: Calculation of relevance for search text block STB(i)] Next, the relevance to the search text block STB(i) is calculated.
[0090] Specifically, by performing a full-text search using the search text block STB(i) as the search criteria, , for each search text block STB(i) of the text block TB to be searched Calculate the relevance.
[0091] Here, for all text blocks TB, the relation to the search text block STB(i) is The relevance may be calculated, and for some sentence blocks TB, the search sentence blocks STB( The relevance for i) may be calculated.
[0092] For example, in the case of patent specifications, we searched for similar documents regarding "background, problem, means, and effect." In this case, you only need to search for the "background, issues, means, and effects" of the document you are searching for. "Embodiment 1" and the like can be excluded from the search.
[0093] In addition, when searching for similar documents in "Embodiment 1," the search target document is searched for in each embodiment. The search target is the state of affairs, and the "background, issues, means, and effects" can be excluded from the search target. Furthermore, if you want to search for similar documents that have an English translation, you can use the Each embodiment of the document that is "present" can be searched.
[0094] In full-text search, the text block TB for calculating the relevance is, for example, an index file. The search text STD is automatically selected based on the information contained in the file. In this case, the text block TB for which the relevance is to be calculated may be specified.
[0095] In this way, the text block to be searched is selected according to the search text block STB(i). By changing the method, the amount of processing can be reduced and the time required for document search can be shortened.
[0096] In example 1 of the document search method, the search text block STB(i) is used as one of the search conditions for the full-text search. As will be described later, the search sentence block STB(i) Each sentence contained in the document may be used as a search criterion for a full-text search (see Example 2 of the document search method). In other words, the number of search conditions is equal to the number of sentences contained in the search sentence block STB(i). Good too.
[0097] There are no particular limitations on the full-text search method, and sequential search, index search, etc. can be used.
[0098] In particular, index search is difficult to perform even when there are many text blocks to search. This is preferable because the speed is less likely to decrease.
[0099] In index search, the text block TB to be searched is scanned in advance, and high Prepare an index file to enable fast searches.
[0100] There is no particular limitation on the method for extracting the character strings that make up the index file. Separating words with spaces), morphological analysis, N-gram (N-character indexing, N (also known as the Gram method) etc. can be used.
[0101] In particular, N-gram is advantageous over morphological analysis in searching for exact matches, and is useful for technical terms, This is preferable because new words, abbreviations, etc. are less likely to cause problems.
[0102] To calculate the relevance, for example, TF-IDF (Term Frequency Inverse Function) It is preferable to use the TF value. The IDF value indicates the frequency of occurrence of each word in a given sentence block. The more a word appears in a block of text, the more likely it is to appear in that block of text. The TF value of the word in the sentence block will be high. It appears in many sentence blocks. The IDF value of a word is small, and the IDF value of a word that appears only in some sentence blocks is high. By calculating the product of the TF score and IDF score of each word, the word characterizes the sentence block. A score can be calculated to determine whether it is a word or not.
[0103] The method for calculating the relevance is not limited to the method using TF-IDF.
[0104] For example, the open source search engine library Apache Lucene It can be used to perform full-text searches.
[0105] FIG. 3B shows an example of calculating the relevance to the search text block STB(1). The first object 110(1) to be searched is the first sentence of each search target document TD. An example is shown for block TB(1).
[0106] [Step A4: Determine second target 120(i) from first target 110(i)] Next, based on the degree of relevance, a second object 120(i) is selected from the first object 110(i). ) to determine
[0107] The number of text blocks TB included in the second object 120(i) is not particularly limited. The object 120(i) is the object for which the similarity is calculated in the next step. The process of calculating the similarity tends to take a long time. The second target 120(i) is determined, and the similarity is calculated after narrowing down the target. This can reduce the time required for searching.
[0108] For example, by sorting the results of the full-text search in step A3 in descending order of relevance, To grasp a text block TB that is highly related to a search text block STB(i). can be done.
[0109] In FIG. 3C, the top 10 most relevant text blocks for the search text block STB(1) are displayed. An example is shown in which the lock TB is used as the second target 120(1). In FIG. 3C, as an example, Sentence block TB4(1) is ranked 1st (Rank 1), sentence block TB1(1) is ranked 2nd ( Rank 2), and sentence block TB9(1) is ranked 10th (Rank 10). This shows the case.
[0110] [Step A5: Calculation of similarity for search text block STB(i)] Next, the similarity to the search sentence block STB(i) is calculated. For each sentence contained in the sentence block STB(i), the sentences contained in the second object 120(i) are The similarity with each of them is calculated.
[0111] In a document retrieval method according to one aspect of the present invention, similarity between sentences is calculated. It is preferable to calculate the similarity based on the degree of agreement of the character "shi" appearance.
[0112] For example, similarity can be calculated using diff, an algorithm that finds differences between documents. This can be done.
[0113] First, as shown in Figure 4A, the first sentence STS1 of the search sentence block STB(1) and The similarity with each sentence included in the second target 120(1) is calculated.
[0114] Next, as shown in Figure 4B, the second sentence STS2 in the search sentence block STB(1) and The similarity with each sentence included in the second target 120(1) is calculated. Classification of each sentence in the chapter block STB(1) with each sentence in the second object 120(1) The similarity is calculated.
[0115] Then, as shown in FIG. 4C, the last sentence STSp(p is an integer greater than or equal to 1), the similarity is calculated based on the search sentence block STB(1). For all sentences included in the second target 120(1), the similarity with each sentence included in the second target 120(1) is calculated. In addition, FIG. 4C shows an example where p is an integer of 3 or more.
[0116] In addition, the calculation of similarity for multiple sentences in the search sentence block STB(1) is performed in parallel. For example, the process shown in FIG. 4A, the process shown in FIG. 4B, and the process shown in FIG. 4C may be performed in a similar manner. The steps may be performed in parallel.
[0117] By using the calculated similarity, sentence blocks similar to the search sentence block STB(1) are found. You can ask for a TB.
[0118] For example, in each sentence block TB, for each sentence in the search sentence block STB(1), The sum of the similarities of the sentences with the highest similarity is calculated, and the sum is used as the search sentence block STB(1). By dividing by the number of sentences in the sentence block TB, the search sentence block STB(1) is It is possible to obtain the normalized similarity.
[0119] In FIG. 5A, in the sentence block TB4(1), the first sentence of the search sentence block STB(1) The sentence with the highest similarity to the second sentence STS1 is the first sentence S1 (similarity is 1), The sentence with the highest similarity to the second sentence STS2 is the second sentence S2 (similarity is 0.9). The sentence with the highest similarity to the last sentence STSp is the third sentence S3 (similarity is 0.5 ) By adding these p similarities and dividing by the number of sentences p, sentence block TB4(1 ) to the search text block STB(1).
[0120] Note that using a value above a threshold for the similarity between sentences can improve search accuracy. For example, when the threshold value is 0.8, the sentence block TB4 shown in FIG. In (1), the similarity of the sentence S3, which has the highest similarity to the last sentence STSp, is 0.5. Therefore, it is not used (considered to be 0) when calculating the sum of similarities.
[0121] [Step A6: Output of results] Then, the text block TB with the highest normalized similarity to the search text block STB(i) is Output.
[0122] FIG. 5B is an example in which text blocks TB (Block) are arranged in descending order of normalized similarity. In addition, an example in which the normalized similarity is expressed as a percentage is shown as the score.
[0123] The full-text search performed in step A3 does not take into account the order of sentences or words, so the calculated relation The degree of similarity is different from the degree of concurrency. By calculating the degree of similarity in step A5, the degree of similarity can be calculated by the method in step A4 (see FIG. 3). C) The 10 sentence blocks TB determined as the second target 120(1) are used as the search sentences. They can be arranged in order of increasing similarity to block STB(1) (Fig. 5B).
[0124] As described above, the search document STD is divided into search sentence blocks STB, and similar sentence blocks are extracted. By searching the lock, similar documents (text blocks) are found for the search text block STB. This allows you to search the entire search document STD as a search criterion. Compared to when the search is performed on a single document or when the search target is the entire document, it is possible to grasp the correspondence between similar parts. This makes it easier to do so.
[0125] In addition, after narrowing down the text blocks to be searched to the first target, then the second target, Since similarity is calculated, the time required for document search can be reduced.
[0126] <Document search method example 2> Next, a modified example of step A3 and subsequent steps will be described with reference to FIGS. When each sentence contained in the sentence block STB(i) is used as a search condition for a full-text search, and explain.
[0127] [Step A3: Calculation of relevance for search text block STB(i)] In step A3 of the example 2 of the document search method, the search text block STB(i) A full-text search is performed using each sentence included in the search target as a search condition. The relevance of each sentence contained in the search sentence block STB(i) is calculated.
[0128] Here, for all text blocks TB, the search text block STB(i) contains The relevance of each sentence may be calculated, and for some sentence blocks TB, a search sentence block may be calculated. The relevance for each sentence included in the lock STB(i) may be calculated.
[0129] By changing the text block to be searched according to the search text block STB(i), This reduces the amount of processing and shortens the time required for document search.
[0130] The full-text search method and the method for calculating relevance should be the same as Example 1 of the document search method. can be done.
[0131] First, as shown in FIG. 6A, the first sentence STS1 in the search sentence block STB(1) is searched. By using the search criteria to perform a full-text search, one of each sentence contained in the first object 110(1) is found. The relevance of the sentence STS1 to the first target 110(1) is calculated. refers to the sentences that make up the multiple sentence blocks TB included in the first object 110(1).
[0132] Next, as shown in Figure 6B, the second sentence STS2 in the search sentence block STB(1) is searched. By using the search criteria to perform a full-text search, two sentences of each sentence included in the first object 110(1) are retrieved. Similarly, the relevance of the search sentence block STB(1) is calculated. Calculate the relevance for each sentence.
[0133] Then, as shown in FIG. 6C, the last sentence STSp(p By calculating the relevance up to the first target 110(1), the sentences included in the first target 110(1) are The relevance of each sentence contained in the search sentence block STB(1) is calculated. FIG. 6C shows an example where p is an integer of 3 or more.
[0134] In addition, full-text searches using each sentence in the search text block STB(1) as a search condition are performed in parallel. For example, the process shown in FIG. 6A, the process shown in FIG. 6B, and the process shown in FIG. 6C may be performed in the following manner: All may be performed in parallel.
[0135] [Step A4: Determine second target 120(i) from first target 110(i)] Next, for each sentence included in the search sentence block STB(i), based on the degree of relevance, A second object 120(i) is determined from among the sentences contained in the first object 110(i).
[0136] The number of sentences included in the second object 120(i) is not particularly limited. ) will be the target for calculating similarity in the next step. The process of selecting the second object 12 from the first object 110(i) tends to take a long time. By determining 0(i) and calculating the similarity after narrowing down the target, the time required for document search is reduced. The time can be shortened.
[0137] For example, by sorting the results of the full-text search in step A3 in descending order of relevance, Identifying sentences that are highly related to each sentence included in the search sentence block STB(i) can be done.
[0138] In FIG. 7A, the high relevance of the first sentence STS1 in the search sentence block STB(1) is An example is shown in Figure 7, where the top 300 sentences are used as the second target 120(1) (STS1). In A, for example, the first sentence of sentence block TB4(1)_S1 is ranked first. (Rank 1), the first sentence of sentence block TB3(1)_S1 is ranked second ( Rank 2), and the sixth sentence of sentence block TB6(1) is TB6(1)_S6. This shows the case where the ranking is 300 (Rank 300).
[0139] In FIG. 7B, the high relevance of the second sentence STS2 in the search sentence block STB(1) is An example is shown in Figure 7, where the top 300 sentences are used as the second target 120(1) (STS2). In B, for example, the second sentence in sentence block TB1(1)_S2 is ranked first. (Rank 1), the second sentence of sentence block TB3(1)_S2 is ranked second ( Rank 2), and the eighth sentence of sentence block TB62(1)_S This shows the case where 8 is ranked 300.
[0140] Then, as shown in FIG. 7C, for the last sentence STSp of the search sentence block STB(1), The second target 120(1) (STSp) is determined as the top 300 most relevant sentences. In FIG. 7C, as an example, the ninth sentence TB2(1)_ S9 is ranked 1 (Rank 1), the eighth sentence in sentence block TB6(1)_S 8 is ranked 2nd (Rank 2), and the 12th sentence of sentence block TB7(1) is 1)_S12 is ranked 300th. For all sentences included in the sentence block STB(1), 1) is determined. Similarly, for all sentences included in the search sentence block STB(i), , respectively, based on the degree of relevance, the sentences included in the first object 110(i) are selected. Determine the object 120(i) of 2.
[0141] [Step A5: Calculation of similarity for search text block STB(i)] Next, the similarity to the search sentence block STB(i) is calculated. For each sentence contained in the sentence block STB(i), the sentences contained in the second object 120(i) are The similarity with each of them is calculated.
[0142] The similarity can be calculated using the same method as in Example 1 of the document search method.
[0143] First, as shown in FIG. 8A, the first sentence STS1 of the search sentence block STB(1) and The similarity with each sentence included in the second target 120(1) (STS1) is calculated.
[0144] Next, as shown in FIG. 8B, the second sentence STS2 of the search sentence block STB(1) and The similarity with each sentence included in the second target 120(1) (STS2) is calculated. Then, each sentence in the search sentence block STB(1) and the sentences included in the second target 120(1) are The similarity with each of them is calculated.
[0145] Then, as shown in FIG. 8C, the last sentence STSp of the search sentence block STB(1) By calculating the similarity, all sentences included in the search sentence block STB(1) are Then, the similarity with each of the sentences included in the second target 120(1) is calculated.
[0146] In addition, the calculation of similarity for multiple sentences in the search sentence block STB(1) is performed in parallel. For example, the process shown in FIG. 8A, the process shown in FIG. 8B, and the process shown in FIG. 8C may be performed in a similar manner. The steps may be performed in parallel.
[0147] By using the calculated similarity, sentence blocks similar to the search sentence block STB(1) are found. You can ask for a TB.
[0148] For example, in each sentence block TB, for each sentence in the search sentence block STB(1), The sum of the similarities of the sentences with the highest similarity is calculated, and the sum is used as the search sentence block STB(1). By dividing by the number of sentences in the sentence block TB, the search sentence block STB(1) is It is possible to obtain the normalized similarity.
[0149] In FIG. 9A, in the sentence block TB4(1), the first sentence of the search sentence block STB(1) The sentence with the highest similarity to the second sentence STS1 is the first sentence S1 (similarity is 1), The sentence with the highest similarity to the second sentence STS2 is the second sentence S2 (similarity is 0.90). In this way, by adding the highest similarity for each of the p sentences and dividing by the number of sentences p, , the normalized similarity of the sentence block TB4(1) to the search sentence block STB(1) In the sentence block TB4(1), the 26th sentence S26 However, the similarity to the first sentence STS1 in the search sentence block STB(1) is high (similar Since the similarity score (0.80) of the first sentence S26 is lower than that of the first sentence S1, the similarity score of S26 is not used.
[0150] Note that using a value above a threshold for the similarity between sentences can improve search accuracy. In the sentence block TB9(1) shown in FIG. 9A, The sentence with the highest similarity to the first sentence STS1 in STB(1) is the second sentence S2 ( The similarity is 0.70, and the sentence with the highest similarity to the second sentence STS2 is the first sentence STS1. S1 (similarity is 0.60), and the most similar to the last sentence STSp is 3 The first sentence is S3 (similarity is 0.60). When no threshold is used, the most significant sentence for each of the p sentences is The similarity values of these three sentences are used to calculate the sum of the high similarity scores. If the value is 0.8, the similarity values of these three sentences are below the threshold, so the similarity It will not be used when calculating the sum (it will be considered 0).
[0151] [Step A6: Output of results] Then, the text block TB with the highest normalized similarity to the search text block STB(i) is Output.
[0152] Figure 9B shows an example of arranging sentence blocks TB in descending order of normalized similarity. Here, e is an example of normalized similarity expressed as a percentage.
[0153] In the example 2 of the document search method, for each sentence included in the search text block STB(i), A sentence that becomes the second object 120(i) is determined from the object 110(i). Among the sentences contained in the chapter block TB, the sentences contained in the search sentence block STB(i) Only highly relevant sentences are calculated to determine the similarity with the sentences contained in the search sentence block STB(i). By narrowing down the target by sentence, you can narrow down the target by sentence block. Compared to the case where similar sentences (and sentence blocks) are not missed, This reduces the time required to calculate the similarity. This can prevent the similarity of lock TBs from becoming too high.
[0154] For example, by using Example 2 of Document Search Method, the top 1 results in Example 1 of Document Search Method (Figure 5B) The sentence blocks that did not rank 0, TB7(1), TB3(1), and TB6(1), are in the top 10. It is also possible that the number of
[0155] In document retrieval method example 2, the similarity of the remaining part is extremely low compared to document retrieval method example 1. However, there are sentence blocks with extremely high similarity (for example, sentences that contain perfect matches). The similarity can be calculated to be high.
[0156] <Document search method example 3> Next, similar sentence blocks are sequentially searched for among multiple search sentence blocks STB. In the example 3 of the document search method, all the search text blocks ST For B, an example of searching for similar sentence blocks is shown below, but it is not limited to this. For the search text block STB, similar text blocks may be searched. 1 shows a flowchart of a document retrieval method.
[0157] The process before the search is the same as in Example 1 of the document search method, so the explanation will be omitted. is omitted.
[0158] [Step B1: Creating multiple search text blocks STB(1) to STB(w)] First, the search document STD is divided into multiple search text blocks STB. Here, w (w is an integer of 2 or more) search text blocks (search text block ST This is an example of dividing B(1) into search sentence blocks STB(w). Step B1 is This can be done in the same manner as step A1 shown in FIG. 3A.
[0159] [Step B2: Selection of search text block STB(i) (i=1)] Next, from among the w search sentence blocks STB, the search sentence block STB to be searched is selected. Select (i) (i is an integer between 1 and w).
[0160] In addition, for some or all of the search sentence block STB, similar sentence blocks can be searched. The search order is not particularly limited.
[0161] In Example 3 of the document search method, a search is performed in order starting from the search text block STB(1). Therefore, in step B2, i=1 is selected.
[0162] [Step B3: Calculation of relevance for search text block STB(i)] Next, the relevance to the search text block STB(i) is calculated.
[0163] Since i=1 was selected in step B2, the first step B3 is The relevance to STB(1) is calculated. The first step B3 is the step shown in FIG. This can be done in the same way as step A3.
[0164] [Step B4: Determine second target 120(i) from among first targets 110(i)] Next, based on the degree of relevance, a second object 120(i) is selected from the first object 110(i). ) to determine
[0165] Since i=1 was selected in step B2, in the first step B4, the Then, the second target 120(1) is determined from the first target 110(1). Step B4 can be performed in the same manner as step A4 shown in FIG. 3C.
[0166] [Step B5: Calculation of similarity for search text block STB(i)] Next, the similarity to the search sentence block STB(i) is calculated. For each sentence contained in the sentence block STB(i), the sentences contained in the second object 120(i) are The similarity with each of them is calculated.
[0167] Since i=1 was selected in step B2, the first step B5 is The similarity to STB(1) is calculated in the first step B5. This can be done in the same manner as step A5 shown in FIG. 5A.
[0168] [Step B6: Has the similarity been calculated for all search sentence blocks STB? (i=w ?)] The above steps B3 to B5 are performed on all search text blocks STB. If there is a search text block STB for which the similarity has not been calculated, Return to step B3 via step B7. Then, for all search text blocks STB, If the similarity has been calculated, the process proceeds to step B8.
[0169] [Step B7: Add 1 to i (i=i+1)] When returning from step B6 to step B3, step B7 adds 1 to i. The second steps B3 to B5 are performed on the search text block STB(2). Step B Repeat steps 3 to B5.
[0170] [Step B8: Output of results] Then, the text block TB with the highest normalized similarity to each search text block STB is extracted. To exert effort.
[0171] Figure 12 shows the sentence blocks TB sorted in descending order of normalized similarity for each search sentence block STB. Furthermore, a value indicating the degree of similarity is output, such as the score shown in Figure 5B. You may also use force.
[0172] As described above, similar sentence blocks were searched for in order for each search sentence block STB. Then, by outputting all the results, the search text block STB of the search document STD is This allows you to search for similar documents (text blocks TB).
[0173] <Document search method example 4> Next, similar sentence blocks are searched in parallel for multiple search sentence blocks STB. In Example 4 of the document search method, all the search text blocks are An example of searching for similar sentence blocks for STB is shown below, but it is not limited to this. A similar sentence block may be searched for using the search sentence block STB shown in FIG. 11. A flowchart of the document retrieval method is shown in FIG.
[0174] The process before the search is the same as in Example 1 of the document search method, so the explanation will be omitted. is omitted.
[0175] [Step C1: Creating multiple search text blocks STB] First, the search document STD is divided into multiple search text blocks STB. Here, w (w is an integer of 2 or more) search text blocks (search text block ST This shows an example of dividing B(1) into search sentence blocks STB(w). Step C1 is This can be done in the same manner as step A1 shown in FIG. 3A.
[0176] The subsequent steps C2 to C5 are performed for two or more search text blocks STB in parallel. In Example 4 of the text search method, for w search text blocks STB, Here is an example of doing this in parallel.
[0177] [Step C2(i): Selection of search text block STB(i)] Next, from among the w search sentence blocks STB, the search sentence block STB to be searched is selected. Select (i) (i is an integer between 1 and w).
[0178] In step C2(1) shown in FIG. 11, i=1 is selected. In parallel with step C2(1), In step C2(2), i=2 is selected, and in step C2(w), i=w Select .
[0179] [Step C3(i): Calculation of relevance for search text block STB(i)] Next, the relevance to the search text block STB(i) is calculated.
[0180] In step C3(1) shown in FIG. 11, the relevance to the search text block STB(1) is calculated. Step C3(1) can be performed in the same manner as step A3 shown in FIG. 3B. do.
[0181] In step C3(2) which is performed in parallel with step C3(1), a search sentence block S The relevance to TB(2) is calculated, and in step C3(w), the search sentence block ST Calculate the relevance to B(w).
[0182] [Step C4(i): Determine the second object 120(i) from the first object 110(i) ] Next, based on the degree of relevance, a second object 120(i) is selected from the first object 110(i). ) to determine
[0183] In step C4(1) shown in FIG. 11, the first target 110 (1 ) to determine the second target 120(1). Step C4(1) is the same as the step shown in FIG. This can be done in the same way as step A4.
[0184] In step C4(2), which is performed in parallel with step C4(1), the search results are sorted based on the degree of relevance. Then, the second object 120(2) is determined from the first object 110(2), and step C4 ( In w), a second object 120 is selected from the first object 110(w) based on the degree of relevance. Determine (w).
[0185] [Step C5: Calculation of similarity for search text block STB(i)] Next, the similarity to the search sentence block STB(i) is calculated. For each sentence contained in the sentence block STB(i), the sentences contained in the second object 120(i) are The similarity with each of them is calculated.
[0186] In step C5(1) shown in FIG. 11, the similarity to the search sentence block STB(1) is calculated. Step C5(1) is the same as step A5 shown in FIGS. 4A to 4C and 5A. This can be done in the same way.
[0187] In step C5(2) which is performed in parallel with step C5(1), a search sentence block S In step C4(w), the similarity to TB(2) is calculated. Calculate the similarity to B(w).
[0188] [Step C6: Output of results] Then, the text block TB with the highest normalized similarity to each search text block STB is extracted. To exert effort.
[0189] Figure 12 shows the sentence blocks TB sorted in descending order of normalized similarity for each search sentence block STB. This is an example of arranging them. Note that the value indicating the degree of similarity is output, such as the score shown in Figure 5B. You may do so.
[0190] As described above, after searching for sentence blocks similar to each search sentence block STB in parallel, By outputting all the results, for each search text block STB of the search document STD, This allows you to search for similar documents (text blocks TB).
[0191] As described above, in the document search method of this embodiment, a sentence block similar to a search sentence block is searched. By searching for locks, it is possible to find parts of the target document that are similar to specific parts of the search document. This allows for accurate search. It is easier to understand the correspondence between similar parts than when searching the entire document. This becomes:
[0192] In addition, in the document search method of this embodiment, the full-text search results are used to search sentence blocks. This narrows down the targets for which similarity is calculated, thereby shortening the time required for document search. This can be done.
[0193] This embodiment mode can be combined with other embodiment modes as appropriate. In the case where multiple configuration examples are shown in one embodiment, the configuration examples may be combined as appropriate. It is possible to do this.
[0194] (Embodiment 2) In this embodiment, a document search system according to one embodiment of the present invention will be described with reference to FIGS. 13 and 14. I will explain.
[0195] The document retrieval system of this embodiment uses the document retrieval method shown in the first embodiment to retrieve documents. Specifically, you can search for a pre-prepared block of text. Search for documents (text blocks) similar to the input search document (search text block) You can search for it.
[0196] <Document search system configuration example 1> FIG. 13 shows a block diagram of the document retrieval system 100. So, let's classify the components by function and show a block diagram as independent blocks. However, it is difficult to completely separate the components into functions, and one component It may involve multiple functions. Also, one function may involve multiple components. For example, the processing performed by the processing unit 103 may be executed by a different server depending on the processing. This may happen.
[0197] The document retrieval system 100 includes at least a processing unit 103. The document retrieval system shown in FIG. The system 100 further includes an input unit 101, a transmission path 102, a storage unit 105, a database 107 and an output unit 109.
[0198] [Input section 101] A search document STD is supplied to the input unit 101 from outside the document search system 100 . The search document STD supplied to the input unit 101 is transmitted via a transmission path 102 to a processing unit 103, The data is supplied to the storage unit 105 or the database 107 .
[0199] [Transmission path 102] The transmission path 102 has a function of transmitting various data. Data is transmitted and received between the memory unit 105, the database 107, and the output unit 109 via a transmission line 1. 02. For example, search document STD, search text block STB The data such as the search target document TD and the text block TB are transmitted via a transmission path 102. is received.
[0200] [Processing section 103] The processing unit 103 receives the data supplied from the input unit 101, the storage unit 105, the database 107, etc. The processing unit 103 has a function of performing calculations using the data. The processing unit 103 stores the calculation results in the storage unit 105. , can be supplied to the database 107, the output unit 109, etc.
[0201] The processing section 103 may use a transistor having a metal oxide in a channel forming region. Since the off-state current of the transistor is extremely low, the transistor can be preferably used as a memory element. It is used as a switch to hold the charge (data) that has flowed into the capacitance element that functions as a This allows data to be retained for a long period of time. By using it for at least one of the register and cache memory of 103, The processing unit 103 is operated only when necessary, and in other cases, the information of the immediately preceding processing is stored in the storage element. By retracting the power supply, the processing unit 103 can be turned off. This makes it possible to achieve low power consumption in document search systems. do.
[0202] In this specification and the like, when an oxide semiconductor or a metal oxide is used for a channel formation region, The transistor is called an Oxide Semiconductor transistor, or OST. The channel formation region of an OS transistor often contains a metal oxide. preferable.
[0203] In this specification, metal oxide refers to a metal oxide in a broad sense. Metal oxides are oxides. Metal oxides are oxide insulators and oxide conductors (including transparent oxide conductors). , oxide semiconductors (also called "OS"), For example, when a metal oxide is used in the semiconductor layer of a transistor, the metal Oxides are sometimes called oxide semiconductors. In other words, metal oxides have amplifying and rectifying properties. and a switching action, the metal oxide is A semiconductor (metal oxide semiconductor), abbreviated as OS. This can be done.
[0204] The metal oxide contained in the channel formation region preferably contains indium (In). When the metal oxide in the panel formation region contains indium, The carrier mobility (electron mobility) of the semiconductor is increased. The oxide is preferably an oxide semiconductor containing element M. Element M is aluminum (Al). Preferably, the element M is gallium (Ga), or tin (Sn). The elements include boron (B), silicon (Si), titanium (Ti), iron (Fe), and nickel. Ni, Germanium (Ge), Yttrium (Y), Zirconium (Zr), Mo Molybdenum (Mo), Lanthanum (La), Cerium (Ce), Neodymium (Nd), Hafnium Examples of elements include Hf, Ta, and W. In some cases, a combination of the above elements may be used. The element M may be, for example, a bond with oxygen. For example, the bond energy with oxygen is higher than that with indium. In addition, the metal oxide of the channel formation region contains zinc (Zn). It is preferable that the metal oxide containing zinc is easily crystallized in some cases.
[0205] The metal oxide contained in the channel formation region is not limited to a metal oxide containing indium. The semiconductor layer may be an indium-free material such as zinc tin oxide or gallium tin oxide. metal oxides containing zinc, metal oxides containing gallium, metal oxides containing tin, etc. It's okay.
[0206] The processing section 103 may also use a transistor containing silicon in the channel forming region. stomach.
[0207] The processing section 103 includes a transistor including an oxide semiconductor in a channel formation region and a It is preferable to use a transistor including silicon in the channel forming region in combination with the transistor.
[0208] The processing unit 103 is, for example, an arithmetic circuit or a central processing unit (CPU). It also has a Recessing Unit.
[0209] The processing unit 103 includes a DSP (Digital Signal Processor), a GP U (Graphics Processing Unit) and other microprocessors The microprocessor may be a Field Programmable Gate Array (FPGA). Gate Array), FPAA (Field Programmable A PLD (Programmable Logic Device) The processing unit 103 may be configured to be realized by a processor. It performs various data processing and program control by interpreting and executing instructions from various programs. The program that can be executed by the processor can be executed by the processor. The data is stored in at least one of the memory area and the storage unit 105 .
[0210] The processing unit 103 may have a main memory. The main memory may be a volatile memory such as a RAM. The memory includes at least one of a volatile memory such as a read-only memory (ROM) and a non-volatile memory such as a read-only memory (ROM).
[0211] Examples of RAM include DRAM (Dynamic Random Access Memory) mory), SRAM (Static Random Access Memory), etc. is used, and a virtual memory space is allocated and used as a working space for the processing unit 103. The operating system and application programs stored in the storage unit 105 , program modules, program data, and lookup tables, etc. These data, programs, and processes loaded into RAM are Each program module is directly accessed and operated by the processing unit 103.
[0212] The ROM contains a BIOS (Basic Input / Output) It can store the ROM (System) and firmware. SCROM, OTPROM (One Time Programmable Read Only Memory), EPROM (Erasable Programmable EPROM is a type of memory that can be read and written by ultraviolet light. UV-EPROM (Ultra-Violet) Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmability) e Read Only Memory), flash memory, etc.
[0213] [Storage section 105] The storage unit 105 has a function of storing the program executed by the processing unit 103. The memory unit 105 stores the calculation results generated by the processing unit 103 and the data input to the input unit 101. The device may have a function to store data, etc.
[0214] The storage unit 105 includes at least one of a volatile memory and a non-volatile memory. The unit 105 may include a volatile memory such as a DRAM or an SRAM. The unit 105 is, for example, a ReRAM (Resistive Random Access Memory). Memory, also known as resistive memory), PRAM (Phase change RAM) andom Access Memory), FeRAM (Ferroelectric Random Access Memory), MRAM (Magnetoresis tive Random Access Memory, also known as magnetoresistive memory), Alternatively, the storage unit 105 may have a nonvolatile memory such as a flash memory. Hard Disk Drive (HDD) and solid Recording media drives such as solid state drives (SSDs) It may have a live performance.
[0215] [Database 107] The database 107 contains at least data such as search target documents TD and text blocks TB. The database 107 also has a function to store the calculation data generated by the processing unit 103. The input unit 101 may have a function of storing the results and data input to the input unit 101. The storage unit 105 and the database 107 do not have to be separated from each other. The document search system includes a storage unit 105 and a database 107. It may have a unit.
[0216] The processing unit 103, the storage unit 105, and the database 107 each have a memory. This can be said to be an example of a non-transitory computer-readable storage medium.
[0217] [Output section 109] The output unit 109 has a function of supplying data to the outside of the document search system 100. For example, the calculation results in the processing unit 103 can be supplied to the outside.
[0218] <Document search system configuration example 2> FIG. 14 shows a block diagram of the document retrieval system 150. The document retrieval system 150 includes The system includes a server 151 and a terminal 152 (such as a personal computer).
[0219] The server 151 includes a communication unit 161a, a transmission path 162, a processing unit 163a, and a database 1 67. Although not shown in FIG. 14, the server 151 further includes a storage unit, an input / output unit, etc. It may have, etc.
[0220] The terminal 152 includes a communication unit 161b, a transmission path 168, a processing unit 163b, a storage unit 165, and an input The terminal 152 further includes a database 169. and the like.
[0221] A user of the document search system 150 inputs a search document STD from a terminal 152 to the server 15. 1. The search document STD is sent from the communication unit 161b to the communication unit 161a.
[0222] The search document STD received by the communication unit 161a is transmitted to the database 1 via the transmission path 162. 67 or a storage unit (not shown). Alternatively, the search document STD is stored in the communication unit 1 61a may be supplied directly to the processing unit 163a.
[0223] The creation of the search sentence block STB, the calculation of the relevance, and the similarity, which have been described in the first embodiment, The calculation of each of the above requires high processing power. The processing capacity of the processing unit 163b of the terminal 152 is higher than that of the processing unit 163b of the terminal 152. Each of the processes is preferably performed in processing unit 163a.
[0224] Then, the processing unit 163a generates search results. The search results are transmitted via the transmission path 162. The search results are stored in the database 167 or a storage unit (not shown). The processing unit 163a may directly supply the information to the communication unit 161a. The search result is output from the communication unit 161a to the terminal 152. It will be sent to 161b.
[0225] [I / O section 169] Data is supplied to the input / output unit 169 from outside the document search system 150. 169 has a function of supplying data to the outside of the document search system 150. As in the search system 100, the input section and the output section may be separate.
[0226] [Transmission path 162 and transmission path 168] The transmission path 162 and the transmission path 168 have a function of transmitting data. Data is transmitted and received between the management unit 163a and the database 167 via a transmission path 162. The communication unit 161b, the processing unit 163b, the storage unit 165, and the input / output unit 16 Data transmission and reception between the devices 9 can be performed via a transmission path 168.
[0227] [Processing Unit 163a and Processing Unit 163b] The processing unit 163a receives data from the communication unit 161a and the database 167. The processing unit 163b has a function of performing calculations using the communication unit 161b, the storage unit 165, It also has the function of performing calculations using data supplied from the input / output unit 169, etc. The processing unit 163a and the processing unit 163b can refer to the description of the processing unit 103. It is preferable that the processing capacity of the processor 163b is higher than that of the processor 163b.
[0228] [Storage section 165] The storage unit 165 has a function of storing the program executed by the processing unit 163b. The storage unit 165 stores the calculation results generated by the processing unit 163b and the data input to the communication unit 161b. The input / output unit 169 has a function of storing data input thereto, and data input to the input / output unit 169.
[0229] [Database 167] The database 167 has a function of storing search target documents TD and text blocks TB. The database 167 also stores the calculation results generated by the processing unit 163a and the Alternatively, the server 151 may have a function to store data input to the server 151. , a storage unit is provided in addition to the database 167, and the storage unit stores the data generated by the processing unit 163a. It may have a function to store the calculation results and data input to the communication unit 161a. stomach.
[0230] [Communication Unit 161a and Communication Unit 161b] The communication units 161a and 161b are used to transmit data between the server 151 and the terminal 152. The communication unit 161a and the communication unit 161b can be a hub, a router, or the like. Data can be sent and received either by wire or wirelessly (e.g. For example, radio waves, infrared rays, etc. may be used.
[0231] This embodiment mode can be combined with other embodiment modes as appropriate. [Explanation of symbols]
[0232] S1: sentence, S2: sentence, S3: sentence, S26: sentence, STB: sentence block for search, STD: search Search document, STS1: sentence, STS2: sentence, STSp: sentence, TB: sentence block, TB1: Text block, TB2: Text block, TB3: Text block, TB4: Text block, TB6: Text block, TB7: Text block, TB9: Text block, TB62: Text Block, TD: Search target document, TD1: Search target document, TD2: Search target document, TDn : Document to be searched, 100: Document search system, 101: Input unit, 102: Transmission path, 103 : processing unit, 105: storage unit, 107: database, 109: output unit, 110: first pair Elephant, 110(i): First Object, 120: Second Object, 120(i): Second Object, 15 0: document search system, 151: server, 152: terminal, 161a: communication unit, 161b: Communication unit, 162: transmission path, 163a: processing unit, 163b: processing unit, 165: storage unit, 16 7: database, 168: transmission line, 169: input / output section
Claims
[Claim 1] A document search system that searches a plurality of search target documents for specific sentence blocks similar to a search document, a first step of dividing the search document to generate w (w is a natural number equal to or greater than 2) search sentence blocks; A second step of determining an i-th (i is a natural number equal to or less than w) search sentence block from among the w search sentence blocks; a third step of calculating the relevance of each of the plurality of text blocks included in the first target to the i-th search text block by performing a full-text search using the i-th search text block as a search condition on a plurality of document blocks that are part of the plurality of search target documents as a first target; a fourth step of determining a second object including a plurality of sentence blocks with high relevance from the first object; a fifth step of calculating a similarity between each sentence in each of the plurality of sentence blocks included in the second target and a sentence included in the i-th search sentence block; a sixth step of repeating the second step to the fifth step for each of the w search document blocks; A document search system having a function of performing a seventh step after the sixth step, of designating a sentence block including a sentence with a high degree of similarity as the specific sentence block.
Citation Information
Patent Citations
Document retrieving method and device therefor
JP1997160928A
Computer program for retrieving relevant document and relevant document retrieving system and method
JP2006092135A
Document index creating device
JP2012104051A
Information processing apparatus, control method, and program
JP2013149259A
Plagiarism Document Detection System Based on Synonym Dictionary and Automatic Reference Citation Mark Attaching System
US20160196342A1