Document Search System

The document search method addresses the challenge of identifying similar document blocks by dividing documents into blocks and calculating relevance and similarity, improving search efficiency and accuracy, particularly for intellectual property documents.

JP7705518B2Active Publication Date: 2025-07-09SEMICON ENERGY LAB CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024089845
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-11-30
Filing Date
2024-06-03
Publication Date
2025-07-09
Estimated Expiration
2039-11-19

AI Technical Summary

Technical Problem

Existing document search methods struggle to accurately identify similar documents at the block level, particularly in the context of creating new specifications, where whole-document similarity calculations can be misleading and time-consuming, and there is a need for precise identification of referenced parts across multiple documents.

Method used

A document search method that divides documents into blocks, calculates relevance and similarity at the block level, using full-text searches and threshold-based similarity calculations to identify specific text or sentence blocks that are similar to a search query, reducing the computational burden and improving accuracy.

Benefits of technology

Enables efficient and accurate identification of similar document blocks, facilitating quicker and more precise searches, especially in intellectual property-related documents, by narrowing down targets based on relevance and similarity thresholds, thus enhancing the efficiency and accuracy of document retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007705518000001
    Figure 0007705518000001
  • Figure 0007705518000002
    Figure 0007705518000002
  • Figure 0007705518000003
    Figure 0007705518000003
Patent Text Reader

Abstract

To retrieve similar documents with high accuracy, for each document block.SOLUTION: A document retrieval system is configured to: retrieve a specific sentence block, from among a plurality of sentence blocks formed by dividing each of multiple search target documents; prepare a first search sentence block, which is a part of a search document; perform full-text search, using the first search sentence block as a search condition, on at least part of the multiple sentence blocks, as a first target, to calculate a first relevance of each of the sentence blocks included in the first target with respect to the first search sentence block; determine a second target from among the first target on the basis of a degree of the first relevance; calculate a first similarity with each of sentences included in the second target, for each of sentences included in the first search sentence block; and retrieve at least one sentence block similar to the first search sentence block, using the first similarity.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One aspect of the present invention relates to a document search method, a document search system, a program, and a non-transitory computer-readable storage medium.

[0002] Note that one aspect of the present invention is not limited to the above technical field. Examples of the technical field of one aspect of the present invention include semiconductor devices, display devices, light-emitting devices, power storage devices, storage devices, electronic devices, lighting devices, input devices (for example, touch sensors, etc.), input / output devices (for example, touch panels, etc.), their driving methods, or their manufacturing methods.

Background Art

[0003] Document search technologies for efficiently searching for target documents from a large number of documents have been actively developed. For example, Patent Document 1 discloses a similar document search method.

[0004] Similar documents may be similar as a whole to the target document, or may have extremely high similarity in some parts and extremely low similarity in other parts.

[0005] In Patent Document 1, the detail level is calculated as an index for determining whether similar documents are similar as a whole or only in part to the target document.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0007] In patent application operations, when creating a new specification (subsequent application specification), the descriptions in the specifications created in the past by the company (prior application specifications) may be referred to or cited. Here, if a translated text of the prior application specification has been created, when creating a translated text of the subsequent application specification, the translated text of the prior application specification can be referred to or cited, and the time required for translating the subsequent application specification can be shortened.

[0008] Depending on the method of searching for similar documents, among the documents with a high similarity calculated for the target document, there may be documents that are not actually similar but have a certain degree of similarity overall, so that documents with a high similarity calculated for the entire document are included. On the other hand, even if the similarity of the remaining parts is extremely low, a document having a part with extremely high similarity (for example, including exactly matching sentences) may have a low similarity calculated for the entire document. For example, for reference or citation of a translated text, the latter document is more preferable than the former document.

[0009] Also, by searching for each sentence of the text, it is possible to find exactly matching sentences, but the flow of the text may be disrupted, and the translation terms may not be unified depending on the specification. Therefore, it is desirable to be able to grasp similar parts in units of sentences including multiple sentences, such as by chapter.

[0010] Also, the specification to be referred to when creating a new specification is not necessarily limited to one. Therefore, it is desirable to be able to easily grasp not only which specification was referred to when creating the new specification, but also which part of which specification was referred to and which part of the new specification was created. And this is true not only for specifications but also for all documents. However, when creating a new document, it is time-consuming and cumbersome to record in detail which part of which document was referred to.

[0011] One aspect of the present invention is to provide a document search method that can search for similar documents for each block of a document. Or, one aspect of the present invention is to provide a document search system that can search for similar documents for each block of a document. Or, one aspect of the present invention is to provide a document search method that can search for similar documents for each block of a document with a simple input method.

[0012] One aspect of the present invention is to provide a document search method that can search for documents with high accuracy. Or, one aspect of the present invention is to provide a document search system that can search for documents with high accuracy. Or, one aspect of the present invention is to realize a document search with high accuracy, particularly a search for documents related to intellectual property, with a simple input method.

[0013] Note that the description of these problems does not prevent the existence of other problems. One aspect of the present invention does not necessarily need to solve all of these problems. It is possible to extract other problems from the description of the specification, drawings, and claims.

Means for Solving the Problems

[0014] One aspect of the present invention is a document search method for searching for a specific text block from a plurality of text blocks created by dividing a plurality of documents to be searched respectively, comprising: preparing a first search text block which is a part of the search document; using the first search text block as a search condition to perform a full-text search on at least a part of the plurality of text blocks as a first target, thereby calculating a first relevance degree of each text block included in the first target with respect to the first search text block; determining a second target from the first target based on the height of the first relevance degree; calculating a first similarity degree between each text included in the second target and each text included in the first search text block; and using the first similarity degree to search for at least one text block similar to the first search text block.

[0015] It is preferable to create a plurality of search text blocks by splitting the search document. At this time, the first search text block is preferably one of the plurality of search text blocks.

[0016] Furthermore, prepare a second search text block, which is another part of the search document, and perform a full-text search using the second search text block as a search condition for at least a part of the plurality of text blocks as a third target, thereby calculating the second relevance of each text block included in the third target to the second search text block. Based on the degree of the second relevance, determine a fourth target from among the third targets, calculate the second similarity between each sentence included in the fourth target and each sentence included in the second search text block, and preferably search for at least one text block similar to the second search text block using the second similarity. At this time, the first target and the third target may be the same or different from each other.

[0017] It is preferable to search for at least one text block similar to the first search text block using a value equal to or greater than a threshold among the first similarities.

[0018] One aspect of the present invention is a document search method for searching for similar text blocks from among a plurality of text blocks created by splitting a plurality of search target documents for each of the plurality of search text blocks, the method including: creating a plurality of search text blocks by splitting a search document; for each of the plurality of search text blocks, performing a full-text search using the search text block as a search condition for at least a part of the plurality of text blocks as a first target, thereby calculating the relevance of each text block included in the first target to the search text block; determining a second target from among the first targets based on the degree of relevance; calculating the similarity between each sentence included in the search text block and each sentence included in the second target for each sentence included in the search text block; and searching for at least one text block similar to the search text block using the similarity.

[0019] One aspect of the present invention is a document search method for searching for a specific sentence block from a plurality of sentence blocks created by dividing a plurality of documents to be searched. A first search sentence block, which is a part of the search document, is prepared, and at least a part of the plurality of sentence blocks is used as a first target. By performing a full-text search using each sentence included in the first search sentence block as a search condition, a first relevance of each sentence included in the first target to each sentence included in the first search sentence block is calculated. For each sentence included in the first search sentence block, a second target is determined from the sentences included in the first target based on the height of the first relevance. For each sentence included in the first search sentence block, a first similarity between each sentence included in the second target and the sentence is calculated, and at least one sentence block similar to the first search sentence block is searched using the first similarity. This is a document search method.

[0020] Preferably, a plurality of search sentence blocks are created by dividing the search document. At this time, the first search sentence block is preferably one of the plurality of search sentence blocks.

[0021] Furthermore, a second search sentence block, which is another part of the search document, is prepared, and at least a part of the plurality of sentence blocks is used as a third target. By performing a full-text search using each sentence included in the second search sentence block as a search condition, a second relevance of each sentence included in the third target to each sentence included in the second search sentence block is calculated. For each sentence included in the second search sentence block, a fourth target is determined from the sentences included in the third target based on the height of the second relevance. For each sentence included in the second search sentence block, a second similarity between each sentence included in the fourth target and the sentence is calculated, and it is preferable to search for at least one sentence block similar to the second search sentence block using the second similarity. At this time, the first target and the third target may be the same or different from each other.

[0022] It is preferable to search for at least one text block similar to the first search text block using a value equal to or greater than the threshold among the first similarities.

[0023] One aspect of the present invention is a document search method for searching for similar text blocks from among a plurality of text blocks created by dividing a plurality of target documents for each of a plurality of search text blocks, the method including: creating a plurality of search text blocks by dividing a search document; for each of the plurality of search text blocks, using at least a part of the plurality of text blocks as a first target, performing a full-text search using each sentence included in the search text block as a search condition, and calculating the degree of relevance of each sentence included in the first target to each sentence included in the search text block; for each sentence included in the search text block, determining a second target from among the sentences included in the first target based on the degree of relevance; for each sentence included in the search text block, calculating the degree of similarity between each sentence included in the second target; and using the degree of similarity to search for at least one text block similar to the search text block.

[0024] One aspect of the present invention is a document search system having a function of performing any of the above document search methods.

[0025] One aspect of the present invention is a document retrieval system that retrieves a specific sentence block from a plurality of sentence blocks created by dividing a plurality of documents to be retrieved, and has a processing unit. The processing unit has a function of preparing a first retrieval sentence block, which is one of a plurality of retrieval sentence blocks created by dividing a retrieval document; a function of calculating a first relevance degree of each sentence block included in a first target with respect to the first retrieval sentence block by performing a full-text search using the first retrieval sentence block as a retrieval condition for at least a part of the plurality of sentence blocks as the first target; a function of determining a second target from the first target based on the height of the first relevance degree; a function of calculating a first similarity degree between each sentence included in the second target and each sentence included in the first retrieval sentence block for each sentence included in the first retrieval sentence block; and a function of retrieving at least one sentence block similar to the first retrieval sentence block using the first similarity degree.

[0026] One aspect of the present invention is a program having a function of causing a processor to execute any of the above document retrieval methods. One aspect of the present invention is a non-transitory computer-readable storage medium storing the program.

[0027] The program may be supplied to a computer by various types of temporary computer-readable storage media. Temporary computer-readable storage media include electrical signals, optical signals, and electromagnetic waves. The temporary computer-readable storage media can supply the program to the computer via a wired communication path such as an electric wire and an optical fiber, or a wireless communication path.

[0028] One aspect of the present invention is a program for searching for a specific sentence block from a plurality of sentence blocks created by dividing a plurality of search target documents, the program comprising: a step of preparing a first search sentence block which is one of a plurality of search sentence blocks created by dividing a search document; a step of calculating a first relevance degree of each sentence block included in a first target with respect to the first search sentence block by performing a full-text search on at least a part of the plurality of sentence blocks using the first search sentence block as a search condition; a step of determining a second target from among the first targets based on the height of the first relevance degree; a step of calculating a first similarity degree between each sentence included in the second target and each sentence included in the first search sentence block for each sentence included in the first search sentence block; and a step of searching for at least one sentence block similar to the first search sentence block using the first similarity degree, to be executed by a processor. One aspect of the present invention is a non-transitory computer-readable storage medium storing the program.

[0029] As the non-transitory computer-readable storage medium, various types of tangible storage media can be used. Examples of the non-transitory computer-readable storage medium include volatile memories such as RAM (Random Access Memory), and non-volatile memories such as ROM (Read Only Memory). Other examples include recording media drives such as hard disc drives (HDD) and solid state drives (SSD), magneto-optical discs, CD-ROMs, CD-Rs, and the like.

Advantages of the Invention

[0030] According to one aspect of the present invention, it is possible to provide a document search method capable of searching for similar documents for each block of a document. According to one aspect of the present invention, it is possible to provide a document search system capable of searching for similar documents for each block of a document. According to one aspect of the present invention, it is possible to provide a document search method capable of searching for similar documents for each block of a document with a simple input method.

[0031] According to one aspect of the present invention, a document search method capable of searching documents with high accuracy can be provided. According to one aspect of the present invention, a document search system capable of searching documents with high accuracy can be provided. According to one aspect of the present invention, with a simple input method, it is possible to achieve highly accurate document search, particularly for documents related to intellectual property.

[0032] Note that the description of these effects does not prevent the existence of other effects. One aspect of the present invention does not necessarily have to have all of these effects. It is possible to extract other effects from the description of the specification, drawings, and claims.

Brief Description of the Drawings

[0033]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

[0034] The embodiments will be described in detail with reference to the drawings. However, the present invention is not limited to the following description, and those skilled in the art can easily understand that the form and details thereof can be variously changed without departing from the spirit and scope of the present invention. Therefore, the present invention should not be construed as being limited to the description of the embodiments shown below.

[0035] In the configuration of the invention described below, the same reference numerals are commonly used for the same parts or parts having the same functions among different drawings, and the repeated description thereof will be omitted. In addition, when referring to the same function, the hatch pattern may be the same and may not be particularly labeled.

[0036] In addition, the position, size, range, etc. of each component shown in the drawings may not represent the actual position, size, range, etc. for the sake of simplicity of understanding. For this reason, the disclosed invention is not necessarily limited to the position, size, range, etc. disclosed in the drawings.

[0037] (Embodiment 1) In this embodiment, a document search method according to an aspect of the present invention will be described with reference to FIGS. 1 to 12. Note that the schematic diagram of the data is an example and is not limited thereto.

[0038] One aspect of the present invention is a document search method for searching for a specific sentence block from a plurality of sentence blocks created by dividing a plurality of search target documents respectively.

[0039] First, prepare a first search sentence block, which is a part of the search document.

[0040] For example, the first search text block can be created by extracting a part of the search document. Alternatively, the first search text block may be one of a plurality of search text blocks created by splitting the search document.

[0041] In the document search method according to one aspect of the present invention, in advance, a plurality of text blocks are created from a plurality of documents to be searched, and further, at the time of search, a search text block is created from the search document. Thereby, a text block similar to the search text block can be searched. Therefore, it becomes easier to grasp the correspondence relationship of similar parts compared to the case where the entire search document is used as a search condition or the case where the search target is the entire document.

[0042] Next, using at least a part of the plurality of text blocks as a first target and performing a full-text search using the first search text block as a search condition, the first relevance of each text block included in the first target to the first search text block is calculated.

[0043] The larger the number of documents to be searched, the larger the number of text blocks. In one aspect of the present invention, for each search text block, the text blocks (first target) to be searched can be narrowed down, so that the processing amount can be reduced and the search speed can be increased.

[0044] Next, a second target is determined from among the first targets based on the degree of the first relevance.

[0045] In full-text search, since the order of sentences and words is not considered, the calculated relevance is different from the similarity. On the other hand, text blocks having words in common with the search text block have a high relevance value, and text blocks with low similarity have a low relevance value. Therefore, the target for calculating the similarity can be narrowed down with high accuracy.

[0046] Next, for each sentence included in the first search text block, the first similarity between each sentence included in the second target is calculated.

[0047] Compared with full-text search, the process of calculating similarity is likely to take a long time. In one aspect of the present invention, the second target is determined from among the first targets, and after narrowing down the targets, the similarity is calculated, so that the time required for document search can be shortened.

[0048] The similarity can be calculated based on the degree of literal match between sentences. Different from full-text search, in the calculation of similarity, the order of words in the sentence is taken into account. Therefore, even if a sentence has words in common with the sentences in the first search sentence block but the word order is different, the similarity will be low.

[0049] Then, using the first similarity, at least one sentence block similar to the first search sentence block is searched for.

[0050] As described above, by using the document search method of one aspect of the present invention, it is possible to easily grasp the description locations of other documents that are similar to specific locations in the search document.

[0051] In addition, the document search method of one aspect of the present invention only needs to input the search document, and since it is not necessary to select keywords for the search, it has the advantages of less burden on the user and less likelihood of differences in search results due to skill.

[0052] In addition, after narrowing down the sentence blocks to be searched in order of the first target and the second target, the similarity is calculated, so that the time required for document search can be shortened.

[0053] In addition, full-text search may be performed by using each sentence included in the first search sentence block as a search condition one by one. In this case, the first relevance of each sentence included in the first target to each sentence included in the first search sentence block is calculated. Then, for each sentence included in the first search sentence block, the second target is determined from among the sentences included in the first target based on the height of the first relevance.

[0054] A document block contains a plurality of sentences. Among the sentences included in the document block, it is not always the case that most of the sentences are similar to those included in the first search document block. Therefore, in order to search for document blocks with a high degree of similarity with high accuracy, it is necessary to calculate the degree of similarity for many document blocks, and the time required to calculate the degree of similarity may become long. In addition, in order to shorten the time required to calculate the degree of similarity, by reducing the number of document blocks that are the second target, there is a risk of missing a document block that contains a sentence with a high degree of similarity.

[0055] Therefore, it is preferable to narrow down the second target from the first target at the sentence level rather than at the document block level. Specifically, for each sentence included in the first search document block, it is preferable to search for sentences with a high degree of relevance and narrow down the targets for calculating the degree of similarity at the sentence level. By narrowing down the targets at the sentence level, it is possible to achieve both suppression of missing sentences (and document blocks) with a high degree of similarity and shortening of the time required to calculate the degree of similarity, as compared with the case of narrowing down the targets at the document block level.

[0056] <Example 1 of Document Search Method> FIG. 1 shows a flowchart of a document search method. As shown in FIG. 1, the document search method according to one aspect of the present invention has six steps of step A1 to step A6.

[0057] In addition, unless otherwise specified, even when explaining a configuration having a plurality of elements (such as a document, a document block, or a sentence), when explaining matters common to each element, variables and alphabets are omitted. For example, when explaining matters common to search target documents TD1, search target documents TD2, and search target documents TDn, etc., it may be described as search target document TD.

[0058] [Preprocessing] First, the processing in the pre-stage of the search will be described with reference to FIG. 2.

[0059] In the preprocessing, a plurality of search target documents TD are divided to create a plurality of document blocks TB.

[0060] In the document search method of this embodiment, a plurality of pre-prepared documents are divided into blocks. And at the time of search, the input search document is also divided into blocks. Thereby, it is possible to search for sentence blocks similar to each block of the search document.

[0061] FIG. 2 shows an example of preparing n (n is an integer of 2 or more) search target documents TD.

[0062] There is no particular limitation on the search target document TD, and various documents can be used.

[0063] Examples of the search target document TD include documents related to intellectual property. Specifically, documents related to intellectual property include specifications used in patent applications, claims, and abstracts. Furthermore, documents related to intellectual property include publications such as patent documents (published patent gazettes, patent gazettes, etc.), utility model gazettes, design gazettes, and papers. It is not limited to publications issued in Japan, and publications issued in various countries around the world can be used as documents related to intellectual property.

[0064] In addition, as the search target document TD, various works including books, papers, reports, columns, or other sentences may be used. Also, medical documents or the like may be used as the search target document TD.

[0065] Also, there is no particular limitation on the language of the document, and for example, documents in Japanese, English, Chinese, Korean, etc. can be used.

[0066] The search target document TD1 shown in FIG. 2 is divided into x (x is an integer of 2 or more) sentence blocks (sentence block TB1(1) to sentence block TB1(x)).

[0067] Also, the search target document TD2 is divided into y (y is an integer of 2 or more) sentence blocks (sentence block TB2(1) to sentence block TB2(y)).

[0068] Further, the document to be searched TDn is divided into z (z is an integer of 2 or more) text blocks (text block TBn(1) to text block TBn(z)).

[0069] For example, when the document to be searched is a document consisting of a plurality of chapters, a plurality of text blocks may be created by dividing each chapter.

[0070] Specifically, in the case of a patent specification, it can be divided into "Background, Problem, Means, and Effect", "Embodiment 1", "Embodiment 2", etc.

[0071] Also, in the case of a paper, it can be divided into "Introduction", "Research Method", "Results", "Discussion", "Conclusion", etc.

[0072] Note that a plurality of text blocks may be created using all the sentences of the document to be searched, or a plurality of text blocks may be created using only the necessary parts of the document to be searched.

[0073] For example, when the document to be searched is a patent specification, a plurality of text blocks may be created without using "Description of Reference Signs".

[0074] The preprocessing is performed at least once before performing document search (before performing step A1). The preprocessing may be performed multiple times according to the application. For example, by performing the preprocessing regularly and adding, updating, or deleting the document to be searched, the search accuracy and convenience can be improved.

[0075] Furthermore, it is preferable to create an index file for full-text search using a plurality of text blocks TB. Thereby, full-text search can be performed in a short time. The configuration of the index file is not particularly limited, and for example, it can have information such as character strings, document names, text block names, and appearance frequencies.

[0076] Further, for example, the index file may have information on whether there is a translated text in each language of the document to be searched TD (or text block TB). Thereby, at the time of search, conditions such as "there is a translated text in English" and "there is a translated text in Chinese" can be specified.

[0077] Next, with reference to FIGS. 3 to 5, the details of the six steps shown in FIG. 1 will be described.

[0078] [Step A1: Creation of a plurality of search text blocks STB] First, a plurality of search text blocks STB are created by dividing the search document STD (FIG. 3A).

[0079] As shown in FIG. 3A, the search document STD is divided into w search text blocks (search text blocks STB(1) to STB(w)), where w is an integer of 2 or more.

[0080] In the document search method of the present embodiment, since the input search document STD is divided into a plurality of search text blocks STB, similar documents (text blocks TB) can be searched for each search text block STB.

[0081] The search document STD is not particularly limited, and various documents can be used.

[0082] Examples of the search document STD include documents related to intellectual property before translation. Thereby, similar translated documents can be searched from the documents to be searched TD, and the translated texts can be referred to or cited.

[0083] In addition, as the search document STD, books, papers, reports, columns, or various works including texts can be used. Thereby, similar documents can be searched from the documents to be searched TD, and it can be confirmed whether there is any suspicion of plagiarism or piracy in the search document STD.

[0084] In addition, a medical document can be used as the search document STD. By searching for medical documents of similar cases using a medical document that describes the progress of treatment, it is possible to refer to the medical treatment and consider what course the patient will take in the future.

[0085] [Step A2: Selection of Search Article Block STB(i)] Next, a search article block STB(i) (where i is an integer from 1 to w) to be searched is selected from among the w search article blocks STB.

[0086] When performing a search for only one search article block STB, the search article block STB may be created by extracting a necessary part from the search document STD in Step A1.

[0087] When performing searches for a plurality of search article blocks STB respectively, they may be searched sequentially one by one (see Example 3 of the document search method), they may be searched in parallel (see Example 4 of the document search method), or they may be searched by combining sequential processing and parallel processing.

[0088] In the document search method of the present embodiment, since a similar article block TB can be searched for each search article block STB, it is possible to accurately and easily grasp the location in the search target document TD that is similar to a specific location in the search document STD.

[0089] [Step A3: Calculation of Relevance for Search Article Block STB(i)] Next, the relevance for the search article block STB(i) is calculated.

[0090] Specifically, by performing a full-text search using the search article block STB(i) as a search condition, the relevance of each article block TB to be searched with respect to the search article block STB(i) is calculated.

[0091] Here, for all the text blocks TB, the relevance to the search text block STB(i) may be calculated, or for some of the text blocks TB, the relevance to the search text block STB(i) may be calculated.

[0092] For example, in the case of a patent specification, when wanting to search for similar documents regarding "background, problems, means, and effects", only the "background, problems, means, and effects" of the document to be searched need to be set as the search target, and "Embodiment 1" etc. can be excluded from the search target.

[0093] Also, when wanting to search for similar documents regarding "Embodiment 1", each embodiment of the document to be searched can be set as the search target, and "background, problems, means, and effects" can be excluded from the search target. Furthermore, when wanting to search for similar documents where "there is an English translation", each embodiment of the document to be searched where "there is an English translation" can be set as the search target.

[0094] In full-text search, the text block TB for which the relevance is calculated is automatically selected, for example, based on the information included in the index file. Or, when inputting the search document STD, the text block TB for which the relevance is calculated may be specified.

[0095] In this way, by changing the text block to be searched according to the search text block STB(i), the processing amount can be reduced and the time required for document search can be shortened.

[0096] Example 1 of the document search method shows the case where the search text block STB(i) is used as one search condition for full-text search. As will be described later, each sentence included in the search text block STB(i) may also be used as a search condition for full-text search (see Example 2 of the document search method). That is, the number of search conditions may be the same as the number of sentences included in the search text block STB(i).

[0097] The full-text search method is not particularly limited, and sequential search, index search, etc. can be used.

[0098] In particular, index search is preferable because the search speed is less likely to decrease even when there are many text blocks TB to be searched.

[0099] In index search, the text blocks TB to be searched are scanned in advance, and an index file that enables high-speed search is prepared.

[0100] There is no particular limitation on the method of extracting the character strings constituting the index file, and word segmentation (separating words with spaces), morphological analysis, N-gram (also called N-character index method, N-gram method, etc.) and the like can be used.

[0101] In particular, N-gram is preferable because it is more advantageous for exact match search than morphological analysis, and technical terms, new words, abbreviations, etc. are less likely to cause problems.

[0102] For calculating the relevance, for example, it is preferable to use TF-IDF (Term Frequency-Inverse Document Frequency). The TF value represents the frequency of occurrence of each word in a text block, and the IDF value represents the degree to which a word appears concentrated in some text blocks. The more a word appears frequently in one text block, the higher the TF value of that word in that text block. The IDF value of a word that appears in many text blocks is small, and the IDF value of a word that appears only in some text blocks is high. By obtaining the product of the TF value and the IDF value of each word, a score can be calculated to determine whether the word is a word that characterizes the text block.

[0103] Note that the calculation of the relevance is not limited to the method using TF-IDF.

[0104] For example, full-text search can be performed using Apache Lucene, which is an open-source search engine library.

[0105] In FIG. 3B, an example of calculating the relevance to the search text block STB(1) is shown. Also shown is an example in which the first target 110(1) to be searched is the first text block TB(1) included in each search target document TD.

[0106] [Step A4: Determine the second target 120(i) from among the first targets 110(i)] Next, based on the degree of relevance, the second target 120(i) is determined from among the first targets 110(i).

[0107] The number of text blocks TB included in the second target 120(i) is not particularly limited. The second target 120(i) will be the target for calculating the similarity in the next step. Compared with full-text search, the process of calculating the similarity is likely to take a longer time. By determining the second target 120(i) from among the first targets 110(i) and narrowing down the target before calculating the similarity, the time required for document search can be shortened.

[0108] For example, by sorting the results of the full-text search in step A3 in descending order of relevance, it is possible to identify the text blocks TB with a high degree of relevance to the search text block STB(i).

[0109] In FIG. 3C, an example of using the top 10 text blocks TB with a high degree of relevance to the search text block STB(1) as the second target 120(1) is shown. In FIG. 3C, as an example, a case where the text block TB4(1) is ranked 1st (Rank 1), the text block TB1(1) is ranked 2nd (Rank 2), and the text block TB9(1) is ranked 10th (Rank 10) is shown.

[0110] [Step A5: Calculate the similarity to the search text block STB(i)] Next, the similarity to the search text block STB(i) is calculated. Specifically, for each sentence included in the search text block STB(i), the similarity to each sentence included in the second target 120(i) is calculated.

[0111] In the document search method according to one aspect of the present invention, the similarity between sentences is determined. Specifically, it is preferable to calculate the similarity based on the literal match degree between sentences.

[0112] For example, the similarity can be calculated using diff, which is an algorithm for obtaining the difference between documents.

[0113] First, as shown in FIG. 4A, the similarity between the first sentence STS1 of the search sentence block STB(1) and each sentence included in the second target 120(1) is calculated.

[0114] Next, as shown in FIG. 4B, the similarity between the second sentence STS2 of the search sentence block STB(1) and each sentence included in the second target 120(1) is calculated. Similarly, the similarity between each sentence of the search sentence block STB(1) and each sentence included in the second target 120(1) is calculated.

[0115] Then, as shown in FIG. 4C, by calculating the similarity up to the last sentence STSp (p is an integer of 1 or more) of the search sentence block STB(1), for all sentences included in the search sentence block STB(1), the similarity with each sentence included in the second target 120(1) is calculated. Note that in FIG. 4C, an example where p is an integer of 3 or more is shown.

[0116] Note that the calculation of the similarity for a plurality of sentences in the search sentence block STB(1) may be performed in parallel. For example, the processes shown in FIG. 4A, FIG. 4B, and FIG. 4C may all be performed in parallel.

[0117] By using the calculated similarity, a sentence block TB similar to the search sentence block STB(1) can be obtained.

[0118] For example, in each text block TB, the sum of the similarities of the sentences with the highest similarity to each sentence in the search text block STB(1) is calculated, and by dividing this sum by the number of sentences in the search text block STB(1), the normalized similarity of the text block TB to the search text block STB(1) can be obtained.

[0119] In Fig. 5A, in the text block TB4(1), the sentence with the highest similarity to the first sentence STS1 of the search text block STB(1) is the first sentence S1 (similarity is 1), the sentence with the highest similarity to the second sentence STS2 is the second sentence S2 (similarity is 0.9), and the sentence with the highest similarity to the last sentence STSp is the third sentence S3 (similarity is 0.5). By adding these p similarities and dividing by the number of sentences p, the normalized similarity of the text block TB4(1) to the search text block STB(1) can be obtained.

[0120] Note that among the similarities between sentences, it is preferable to use values equal to or greater than the threshold value because the search accuracy can be improved. For example, when the threshold value is 0.8, in the text block TB4(1) shown in Fig. 5A, since the similarity of the sentence S3 with the highest similarity to the last sentence STSp is 0.5, it will not be used (considered as 0) when calculating the sum of similarities.

[0121] [Step A6: Output of Results] Then, the text block TB with a high normalized similarity to the search text block STB(i) is output.

[0122] Fig. 5B is an example in which the text blocks TB (Block) are arranged in descending order of the normalized similarity. Also shown is an example in which the normalized similarity is expressed as a percentage as Score.

[0123] In the full-text search performed in step A3, since the order of sentences and words is not considered, the calculated relevance is different from the similarity. By calculating the similarity in step A5, the 10 sentence blocks TB determined as the second target 120(1) in step A4 (Fig. 3C) can be arranged in descending order of similarity to the search sentence block STB(1) (Fig. 5B).

[0124] As described above, by dividing the search document STD into search sentence blocks STB and searching for similar sentence blocks, similar documents (sentence blocks TB) can be searched for the search sentence block STB. This makes it easier to grasp the correspondence of similar parts compared to the case where the entire search document STD is used as the search condition or when the search target is the entire document.

[0125] Also, since the calculation of similarity is performed after narrowing down the sentence blocks to be searched in order as the first target and the second target, the time required for document search can be shortened.

[0126] <Example 2 of Document Search Method> Next, with reference to FIGS. 6 to 9, a modified example after step A3 will be described. Specifically, the case where each sentence included in the search sentence block STB(i) is used as the search condition for full-text search will be described.

[0127] [Step A3: Calculation of Relevance for Search Sentence Block STB(i)] In step A3 of Example 2 of the document search method, full-text search is performed using each sentence included in the search sentence block STB(i) as the search condition. Thereby, the relevance of each sentence included in the search target to each sentence included in the search sentence block STB(i) is calculated.

[0128] Here, the relevance of each sentence included in the search sentence block STB(i) may be calculated for all the sentence blocks TB, or the relevance of each sentence included in the search sentence block STB(i) may be calculated for some of the sentence blocks TB.

[0129] By changing the text block to be searched according to the search text block STB(i), the processing amount can be reduced, and the time required for document search can be shortened.

[0130] For the full-text search method and the method of calculating relevance, the same method as in Example 1 of the document search method can be used.

[0131] First, as shown in FIG. 6A, by performing a full-text search using the first sentence STS1 of the search text block STB(1) as a search condition, the relevance of each sentence included in the first target 110(1) to the first sentence STS1 is calculated. Note that the sentences included in the first target 110(1) refer to the sentences constituting the plurality of text blocks TB included in the first target 110(1).

[0132] Next, as shown in FIG. 6B, by performing a full-text search using the second sentence STS2 of the search text block STB(1) as a search condition, the relevance of each sentence included in the first target 110(1) to the second sentence STS2 is calculated. Similarly, the relevance of each sentence of the search text block STB(1) is calculated.

[0133] Then, as shown in FIG. 6C, by calculating the relevance up to the last sentence STSp (p is an integer of 2 or more) of the search text block STB(1), the relevance of the sentences included in the first target 110(1) to each sentence included in the search text block STB(1) is calculated. Note that in FIG. 6C, an example where p is an integer of 3 or more is shown.

[0134] Note that the full-text search using each sentence of the search text block STB(1) as a search condition may be performed in parallel. For example, the processes shown in FIG. 6A, the process shown in FIG. 6B, and the process shown in FIG. 6C may all be performed in parallel.

[0135] [Step A4: Determine the second target 120(i) from among the first targets 110(i)] Next, for each sentence included in the search text block STB(i), based on the degree of relevance, the second target 120(i) is determined from among the sentences included in the first target 110(i).

[0136] The number of sentences included in the second target 120(i) is not particularly limited. The second target 120(i) becomes the target for calculating the similarity in the next step. Compared with full-text search, the process of calculating the similarity is likely to take a long time. By determining the second target 120(i) from among the first target 110(i) and narrowing down the target before calculating the similarity, the time required for document search can be shortened.

[0137] For example, by sorting the results of the full-text search in step A3 in descending order of relevance, it is possible to grasp the sentences with high relevance for each sentence included in the search text block STB(i).

[0138] FIG. 7A shows an example in which the top 300 sentences with high relevance to the first sentence STS1 of the search text block STB(1) are used as the second target 120(1)(STS1). In FIG. 7A, as an example, the first sentence TB4(1)_S1 of the text block TB4(1) is in the first place (Rank 1), the first sentence TB3(1)_S1 of the text block TB3(1) is in the second place (Rank 2), and the sixth sentence TB6(1)_S6 of the text block TB6(1) is in the 300th place (Rank 300).

[0139] FIG. 7B shows an example in which the top 300 sentences with high relevance to the second sentence STS2 of the search text block STB(1) are used as the second target 120(1)(STS2). In FIG. 7B, as an example, the second sentence TB1(1)_S2 of the text block TB1(1) is in the first place (Rank 1), the second sentence TB3(1)_S2 of the text block TB3(1) is in the second place (Rank 2), and the eighth sentence TB62(1)_S8 of the text block TB62(1) is in the 300th place (Rank 300).

[0140] Then, as shown in FIG. 7C, the second target 120(1)(STSp) is determined as the top 300 sentences with a high degree of relevance to the last sentence STSp of the search sentence block STB(1). In FIG. 7C, as an example, the ninth sentence TB2(1)_S9 of the sentence block TB2(1) is ranked first (Rank 1), the eighth sentence TB6(1)_S8 of the sentence block TB6(1) is ranked second (Rank 2), and the 12th sentence TB7(1)_S12 of the sentence block TB7(1) is ranked 300th (Rank 300). As described above, for all the sentences included in the search sentence block STB(1), the second target 120(1) is determined respectively. Similarly, for all the sentences included in the search sentence block STB(i), the second target 120(i) is determined respectively from the sentences included in the first target 110(i) based on the degree of relevance.

[0141] [Step A5: Calculation of similarity for the search sentence block STB(i)] Next, the similarity for the search sentence block STB(i) is calculated. Specifically, for each sentence included in the search sentence block STB(i), the similarity with each sentence included in the second target 120(i) is calculated.

[0142] As the method for calculating the similarity, a method similar to Example 1 of the document search method can be used.

[0143] First, as shown in FIG. 8A, the similarity between the first sentence STS1 of the search sentence block STB(1) and each sentence included in the second target 120(1)(STS1) is calculated.

[0144] Next, as shown in FIG. 8B, the similarity between the second sentence STS2 of the search sentence block STB(1) and each sentence included in the second target 120(1)(STS2) is calculated. Similarly, the similarity between each sentence of the search sentence block STB(1) and each sentence included in the second target 120(1) is calculated.

[0145] Then, as shown in FIG. 8C, by calculating the similarity up to the last sentence STSp of the search text block STB(1), the similarity of each sentence included in the search text block STB(1) with each sentence included in the second target 120(1) is calculated.

[0146] Note that the calculation of the similarity for a plurality of sentences in the search text block STB(1) may be performed in parallel. For example, the processes shown in FIGS. 8A, 8B, and 8C may all be performed in parallel.

[0147] By using the calculated similarity, a text block TB similar to the search text block STB(1) can be obtained.

[0148] For example, in each text block TB, the sum of the similarities of the sentences with the highest similarity to each sentence of the search text block STB(1) is calculated, and the sum is divided by the number of sentences in the search text block STB(1) to obtain the normalized similarity of the text block TB with respect to the search text block STB(1).

[0149] In FIG. 9A, in the text block TB4(1), the sentence with the highest similarity to the first sentence STS1 of the search text block STB(1) is the first sentence S1 (similarity is 1), and the sentence with the highest similarity to the second sentence STS2 is the second sentence S2 (similarity is 0.90). In this way, by adding the highest similarities for each of the p sentences and dividing by the number of sentences p, the normalized similarity of the text block TB4(1) with respect to the search text block STB(1) can be obtained. Note that in the text block TB4(1), the 26th sentence S26 also has a high similarity (similarity 0.80) to the first sentence STS1 of the search text block STB(1), but since it is lower than the first sentence S1, the similarity value of S26 is not used.

[0150] Note that, among the similarities between sentences, it is preferable to use a value equal to or greater than the threshold value in order to improve the accuracy of the search. In the sentence block TB9(1) shown in FIG. 9A, the sentence with the highest similarity to the first sentence STS1 of the search sentence block STB(1) is the second sentence S2 (similarity is 0.70), the sentence with the highest similarity to the second sentence STS2 is the first sentence S1 (similarity is 0.60), and the sentence with the highest similarity to the last sentence STSp is the third sentence S3 (similarity is 0.60). When not using the threshold value, the similarity values of these three sentences are used to calculate the sum of the highest similarities for each of the p sentences. On the other hand, for example, when the threshold value is 0.8, since the similarity values of these three sentences are less than the threshold value, they are not used (considered as 0) when calculating the sum of similarities.

[0151] [Step A6: Output of Results] Then, a sentence block TB with a high normalized similarity to the search sentence block STB(i) is output.

[0152] FIG. 9B is an example in which sentence blocks TB are arranged in descending order of normalized similarity. Also shown is an example in which the normalized similarity is expressed as a percentage as Score.

[0153] In Example 2 of the document search method, for each sentence included in the search sentence block STB(i), a sentence that becomes the second target 120(i) is determined from among the first targets 110(i). Therefore, among the sentences included in the sentence block TB, only the sentences with a high relevance to the sentences included in the search sentence block STB(i) can calculate the similarity to the sentences included in the search sentence block STB(i). By narrowing down the target at the sentence level, it is possible to suppress the omission of sentences (and sentence blocks) with high similarity compared to the case of narrowing down the target at the sentence block level, and to shorten the time required for calculating the similarity. Also, it is possible to prevent the similarity of sentence blocks TB that are not actually similar from becoming high.

[0154] For example, by using Example 2 of the document search method, it is possible that the text blocks TB7(1), TB3(1), and TB6(1) that did not rank in the top 10 in Example 1 of the document search method (Figure 5B) will rank in the top 10 (Figure 9B).

[0155] Compared with Example 1 of the document search method, Example 2 of the document search method can calculate a high similarity degree for a text block that has a part with extremely high similarity (for example, includes a text that exactly matches), even if the similarity of the remaining part is extremely low.

[0156] <Example 3 of the document search method> Next, a method for sequentially searching for similar text blocks among a plurality of search text blocks STB will be described. In Example 3 of the document search method, an example of searching for similar text blocks for all search text blocks STB is shown, but it is not limited to this, and similar text blocks may be searched for some of the search text blocks STB. Figure 10 shows a flowchart of the document search method.

[0157] Note that the processing in the previous stage of the search is the same as that in Example 1 of the document search method, so the description is omitted.

[0158] [Step B1: Creation of a plurality of search text blocks STB(1) to STB(w)] First, a plurality of search text blocks STB are created by splitting the search document STD. Here, an example of splitting into w search text blocks (search text block STB(1) to search text block STB(w), where w is an integer of 2 or more) is shown. Step B1 can be performed in the same manner as step A1 shown in Figure 3A.

[0159] [Step B2: Selection of the search text block STB(i) (i = 1)] Next, a search text block STB(i) (i is an integer from 1 to w) to be searched is selected from among the w search text blocks STB.

[0160] Note that, for some or all of the search text blocks STB, the order in which similar text blocks are searched is not particularly limited.

[0161] In Example 3 of the document search method, an example of performing the search in order from the search text block STB(1) is shown. Therefore, in step B2, i = 1 is selected.

[0162] [Step B3: Calculation of relevance for the search text block STB(i)] Next, the relevance for the search text block STB(i) is calculated.

[0163] Since i = 1 is selected in step B2, in the first step B3, the relevance for the search text block STB(1) is calculated. The first step B3 can be performed in the same manner as step A3 shown in FIG. 3B.

[0164] [Step B4: Determine the second target 120(i) from among the first targets 110(i)] Next, based on the degree of relevance, the second target 120(i) is determined from among the first targets 110(i).

[0165] Since i = 1 is selected in step B2, in the first step B4, based on the degree of relevance, the second target 120(1) is determined from among the first targets 110(1). The first step B4 can be performed in the same manner as step A4 shown in FIG. 3C.

[0166] [Step B5: Calculation of similarity for the search text block STB(i)] Next, the similarity for the search text block STB(i) is calculated. Specifically, for each sentence included in the search text block STB(i), the similarity with each sentence included in the second target 120(i) is calculated.

[0167] Since i = 1 is selected in step B2, in the first step B5, the similarity to the search text block STB(1) is calculated. The first step B5 can be performed in the same manner as step A5 shown in FIGS. 4A to 4C and FIG. 5A.

[0168] [Step B6: Have similarities been calculated for all search text blocks STB (i = w?)] The processes from step B3 to step B5 above are sequentially performed for all search text blocks STB. If there is a search text block STB for which the similarity has not been calculated, the process returns to step B3 via step B7. Then, if similarities have been calculated for all search text blocks STB, the process proceeds to step B8.

[0169] [Step B7: Add 1 to i (i = i + 1)] When returning from step B6 to step B3, as step B7, 1 is added to i. That is, the second steps B3 to B5 are performed for the search text block STB(2). In this way, steps B3 to B5 are repeatedly performed until the similarity to the search text block STB(w) is calculated.

[0170] [Step B8: Output of results] Then, the text block TB with a high normalized similarity to each search text block STB is output.

[0171] FIG. 12 is an example in which text blocks TB are arranged in descending order of normalized similarity for each search text block STB. Further, a value indicating the degree of similarity, such as Score shown in FIG. 5B, may be output.

[0172] As described above, for each search text block STB, after sequentially searching for similar text blocks and then outputting all the results, similar documents (text blocks TB) can be searched for each search text block STB of the search document STD.

[0173] <Example 4 of document search method> Next, a method for searching for similar text blocks in parallel for a plurality of search text blocks STB will be described. Note that in Example 4 of the document search method, an example of searching for similar text blocks for all search text blocks STB is shown, but it is not limited to this, and similar text blocks may be searched for some of the search text blocks STB. FIG. 11 shows a flowchart of the document search method.

[0174] Note that the processing in the previous stage of the search is the same as that in Example 1 of the document search method, so the description is omitted.

[0175] [Step C1: Creation of a plurality of search text blocks STB] First, a plurality of search text blocks STB are created by dividing the search document STD. Here, an example of dividing into w (where w is an integer of 2 or more) search text blocks (search text block STB(1) to search text block STB(w)) is shown. Step C1 can be performed in the same manner as step A1 shown in FIG. 3A.

[0176] The processing of subsequent steps C2 to C5 can be performed in parallel for two or more search text blocks STB. In Example 4 of the text search method, an example of performing in parallel for w search text blocks STB is shown.

[0177] [Step C2(i): Selection of search text block STB(i)] Next, a search text block STB(i) (where i is an integer from 1 to w) to be searched is selected from among the w search text blocks STB.

[0178] In step C2(1) shown in FIG. 11, i = 1 is selected. In step C2(2) performed in parallel with step C2(1), i = 2 is selected, and in step C2(w), i = w is selected.

[0179] [Step C3(i): Calculation of relevance for search text block STB(i)] Next, the relevance to the search text block STB(i) is calculated.

[0180] In step C3(1) shown in FIG. 11, the relevance to the search text block STB(1) is calculated. Step C3(1) can be performed in the same manner as step A3 shown in FIG. 3B.

[0181] In step C3(2) performed in parallel with step C3(1), the relevance to the search text block STB(2) is calculated, and in step C3(w), the relevance to the search text block STB(w) is calculated.

[0182] [Step C4(i): Determine the second target 120(i) from among the first targets 110(i)] Next, based on the degree of relevance, the second target 120(i) is determined from among the first targets 110(i).

[0183] In step C4(1) shown in FIG. 11, based on the degree of relevance, the second target 120(1) is determined from among the first targets 110(1). Step C4(1) can be performed in the same manner as step A4 shown in FIG. 3C.

[0184] In step C4(2) performed in parallel with step C4(1), based on the degree of relevance, the second target 120(2) is determined from among the first targets 110(2), and in step C4(w), based on the degree of relevance, the second target 120(w) is determined from among the first targets 110(w).

[0185] [Step C5: Calculate the similarity to the search text block STB(i)] Next, the similarity to the search text block STB(i) is calculated. Specifically, for each sentence included in the search text block STB(i), the similarity to each sentence included in the second target 120(i) is calculated.

[0186] In step C5(1) shown in FIG. 11, the similarity to the search text block STB(1) is calculated. Step C5(1) can be performed in the same manner as step A5 shown in FIGS. 4A to 4C and FIG. 5A.

[0187] In step C5(2) performed in parallel with step C5(1), the similarity to the search text block STB(2) is calculated, and in step C4(w), the similarity to the search text block STB(w) is calculated.

[0188] [Step C6: Output of Results] Then, a text block TB with a high normalized similarity to each search text block STB is output.

[0189] FIG. 12 is an example in which text blocks TB are arranged in descending order of normalized similarity for each search text block STB. Note that a value indicating the degree of similarity may be output as Score shown in FIG. 5B.

[0190] As described above, after searching for text blocks similar to each search text block STB in parallel and outputting all the results, similar documents (text blocks TB) can be searched for each search text block STB of the search document STD.

[0191] As described above, in the document search method of the present embodiment, by searching for text blocks similar to the search text blocks, the description locations of the search target documents similar to specific locations of the search document can be accurately searched. Thereby, it becomes easier to grasp the correspondence relationship of similar locations compared to the case where the entire search document is used as the search condition or the case where the search target is the entire document.

[0192] Also, in the document search method of the present embodiment, the target for calculating the similarity to the search text block is narrowed down using the full-text search results. Thereby, the time related to document search can be shortened.

[0193] This embodiment can be appropriately combined with other embodiments. Also, in this specification, when multiple configuration examples are shown within one embodiment, the configuration examples can be appropriately combined.

[0194] (Embodiment 2) In this embodiment, a document search system according to one aspect of the present invention will be described with reference to FIGS. 13 and 14.

[0195] The document search system of this embodiment can search for documents using the document search method shown in Embodiment 1. Specifically, it is possible to search for documents (document blocks) similar to the input search document (its search document block) by using the pre-prepared document blocks as the search targets.

[0196] <Configuration Example 1 of Document Search System> FIG. 13 shows a block diagram of the document search system 100. In the drawings attached to this specification, the components are classified by function and shown as independent blocks in the block diagram. However, in actuality, it is difficult to completely separate the components by function, and one component may be related to multiple functions. Also, one function may be related to multiple components. For example, the processing performed by the processing unit 103 may be executed on different servers depending on the processing.

[0197] The document search system 100 has at least a processing unit 103. The document search system 100 shown in FIG. 13 further has an input unit 101, a transmission path 102, a storage unit 105, a database 107, and an output unit 109.

[0198] [Input Unit 101] The input unit 101 is supplied with a search document STD from outside the document search system 100. The search document STD supplied to the input unit 101 is supplied to the processing unit 103, the storage unit 105, or the database 107 via the transmission path 102.

[0199] [Transmission Path 102] The transmission path 102 has a function of transmitting various data. Data transmission and reception between the input unit 101, the processing unit 103, the storage unit 105, the database 107, and the output unit 109 can be performed via the transmission path 102. For example, data such as the search document STD, the search text block STB, the document to be searched TD, and the text block TB are transmitted and received via the transmission path 102.

[0200] [Processing unit 103] The processing unit 103 has a function of performing calculations using the data supplied from the input unit 101, the storage unit 105, the database 107, etc. The processing unit 103 can supply the calculation results to the storage unit 105, the database 107, the output unit 109, etc.

[0201] It is preferable to use a transistor having a metal oxide in the channel formation region in the processing unit 103. Since the off-current of the transistor is extremely low, by using it as a switch for holding the charge (data) flowing into the capacitive element that functions as a memory element, the data holding period can be ensured for a long time. By using this characteristic in at least one of the register and the cache memory of the processing unit 103, the processing unit 103 can be operated only when necessary, and in other cases, the processing unit 103 can be turned off by saving the information of the previous processing in the memory element. That is, normally-off computing becomes possible, and the power consumption of the document search system can be reduced.

[0202] In this specification, etc., a transistor using an oxide semiconductor or a metal oxide in the channel formation region is called an Oxide Semiconductor transistor, or an OS transistor. The channel formation region of the OS transistor preferably has a metal oxide.

[0203] In this specification and the like, a metal oxide refers to an oxide of a metal in a broad sense. Metal oxides are classified into oxide insulators, oxide conductors (including transparent oxide conductors), oxide semiconductors (also referred to as Oxide Semiconductor or simply OS), etc. For example, when a metal oxide is used for the semiconductor layer of a transistor, the metal oxide may be referred to as an oxide semiconductor. That is, when a metal oxide has at least one of an amplification action, a rectification action, and a switching action, the metal oxide can be referred to as a metal oxide semiconductor, abbreviated as OS.

[0204] The metal oxide included in the channel formation region preferably contains indium (In). When the metal oxide included in the channel formation region is a metal oxide containing indium, the carrier mobility (electron mobility) of the OS transistor increases. Also, the metal oxide included in the channel formation region is preferably an oxide semiconductor containing element M. Element M is preferably aluminum (Al), gallium (Ga), or tin (Sn). Other elements applicable to element M include boron (B), silicon (Si), titanium (Ti), iron (Fe), nickel (Ni), germanium (Ge), yttrium (Y), zirconium (Zr), molybdenum (Mo), lanthanum (La), cerium (Ce), neodymium (Nd), hafnium (Hf), tantalum (Ta), tungsten (W), etc. However, there may be cases where a plurality of the aforementioned elements are combined as element M. Element M is, for example, an element with a high binding energy with oxygen. For example, it is an element with a higher binding energy with oxygen than indium. Also, the metal oxide included in the channel formation region preferably contains zinc (Zn). A metal oxide containing zinc may be likely to crystallize.

[0205] The metal oxide included in the channel formation region is not limited to a metal oxide containing indium. The semiconductor layer may be, for example, a metal oxide containing zinc but not indium, such as zinc tin oxide or gallium tin oxide, a metal oxide containing gallium, a metal oxide containing tin, or the like.

[0206] Further, in the processing unit 103, a transistor including silicon in the channel formation region may be used.

[0207] Further, in the processing unit 103, it is preferable to use in combination a transistor including an oxide semiconductor in the channel formation region and a transistor including silicon in the channel formation region.

[0208] The processing unit 103 includes, for example, an arithmetic circuit or a central processing unit (CPU).

[0209] The processing unit 103 may include a microprocessor such as a DSP (Digital Signal Processor) or a GPU (Graphics Processing Unit). The microprocessor may be configured to be realized by a programmable logic device (PLD) such as an FPGA (Field Programmable Gate Array) or an FPAA (Field Programmable Analog Array). The processing unit 103 can perform various data processes and program controls by interpreting and executing instructions from various programs by the processor. Programs executable by the processor are stored in at least one of the memory area of the processor and the storage unit 105.

[0210] The processing unit 103 may include a main memory. The main memory includes at least one of a volatile memory such as a RAM and a non-volatile memory such as a ROM.

[0211] As the RAM, for example, DRAM (Dynamic Random Access Memory), SRAM (Static Random Access Memory), etc. are used, and a memory space is virtually allocated and used as the working space of the processing unit 103. The operating system, application programs, program modules, program data, look-up tables, etc. stored in the storage unit 105 are loaded into the RAM for execution. These data, programs, and program modules loaded into the RAM are directly accessed and operated on by the processing unit 103 respectively.

[0212] The ROM can store the BIOS (Basic Input / Output System), firmware, etc. that do not require rewriting. Examples of the ROM include mask ROM, OTPROM (One Time Programmable Read Only Memory), EPROM (Erasable Programmable Read Only Memory), etc. Examples of the EPROM include UV-EPROM (Ultra-Violet Erasable Programmable Read Only Memory) that enables erasure of stored data by ultraviolet irradiation, EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory, etc.

[0213] [Storage Unit 105] The storage unit 105 has a function of storing the programs executed by the processing unit 103. Further, the storage unit 105 may have a function of storing the calculation results generated by the processing unit 103, the data input to the input unit 101, etc.

[0214] The storage unit 105 has at least one of a volatile memory and a non-volatile memory. The storage unit 105 may have, for example, a volatile memory such as DRAM or SRAM. The storage unit 105 may have, for example, a non-volatile memory such as ReRAM (Resistive Random Access Memory, also referred to as a resistive change memory), PRAM (Phase change Random Access Memory), FeRAM (Ferroelectric Random Access Memory), MRAM (Magnetoresistive Random Access Memory, also referred to as a magnetic resistance memory), or a flash memory. Further, the storage unit 105 may have a recording media drive such as a hard disc drive (HDD) and a solid state drive (SSD).

[0215] [Database 107] The database 107 has at least a function of storing data such as the search target document TD and the text block TB. Further, the database 107 may have a function of storing the calculation result generated by the processing unit 103, the data input to the input unit 101, and the like. Note that the storage unit 105 and the database 107 may not be separated from each other. For example, the document search system may have a storage unit having the functions of both the storage unit 105 and the database 107.

[0216] Note that the memories of the processing unit 103, the storage unit 105, and the database 107 can each be regarded as an example of a non-transitory computer-readable storage medium.

[0217] [Output unit 109] The output unit 109 has a function of supplying data to the outside of the document search system 100. For example, the calculation result in the processing unit 103 can be supplied to the outside.

[0218] [Configuration example 2 of document search system] FIG. 14 shows a block diagram of a document search system 150. The document search system 150 includes a server 151 and a terminal 152 (such as a personal computer).

[0219] The server 151 includes a communication unit 161a, a transmission path 162, a processing unit 163a, and a database 167. Although not shown in FIG. 14, the server 151 may further include a storage unit, an input / output unit, etc.

[0220] The terminal 152 includes a communication unit 161b, a transmission path 168, a processing unit 163b, a storage unit 165, and an input / output unit 169. Although not shown in FIG. 14, the terminal 152 may further include a database, etc.

[0221] A user of the document search system 150 inputs a search document STD from the terminal 152 to the server 151. The search document STD is transmitted from the communication unit 161b to the communication unit 161a.

[0222] The search document STD received by the communication unit 161a is stored in the database 167 or a storage unit (not shown) via the transmission path 162. Alternatively, the search document STD may be directly supplied from the communication unit 161a to the processing unit 163a.

[0223] The creation of the search text block STB, the calculation of the relevance, and the calculation of the similarity described in Embodiment 1 each require high processing capabilities. The processing unit 163a of the server 151 has higher processing capabilities than the processing unit 163b of the terminal 152. Therefore, these processes are preferably performed by the processing unit 163a respectively.

[0224] Then, a search result is generated by the processing unit 163a. The search result is stored in the database 167 or a storage unit (not shown) via the transmission path 162. Alternatively, the search result may be directly supplied from the processing unit 163a to the communication unit 161a. Thereafter, the search result is output from the server 151 to the terminal 152. The search result is transmitted from the communication unit 161a to the communication unit 161b.

[0225] [Input / Output Unit 169] Data is supplied to the input / output unit 169 from outside the document search system 150. The input / output unit 169 has a function of supplying data outside the document search system 150. Note that, like the document search system 100, the input unit and the output unit may be separated.

[0226] [Transmission Lines 162 and 168] The transmission lines 162 and 168 have a function of transmitting data. The transmission and reception of data among the communication unit 161a, the processing unit 163a, and the database 167 can be performed via the transmission line 162. The transmission and reception of data among the communication unit 161b, the processing unit 163b, the storage unit 165, and the input / output unit 169 can be performed via the transmission line 168.

[0227] [Processing Units 163a and 163b] The processing unit 163a has a function of performing calculations using the data supplied from the communication unit 161a, the database 167, etc. The processing unit 163b has a function of performing calculations using the data supplied from the communication unit 161b, the storage unit 165, the input / output unit 169, etc. For the description of the processing units 163a and 163b, reference can be made to the description of the processing unit 103. It is preferable that the processing unit 163a has a higher processing capacity than the processing unit 163b.

[0228] [Storage Unit 165] The storage unit 165 has a function of storing the program executed by the processing unit 163b. Also, the storage unit 165 has a function of storing the calculation results generated by the processing unit 163b, the data input to the communication unit 161b, the data input to the input / output unit 169, etc.

[0229] [Database 167] The database 167 has a function of storing the search target document TD and the text block TB. Further, the database 167 may have a function of storing the calculation result generated by the processing unit 163a, the data input to the communication unit 161a, and the like. Alternatively, the server 151 may have a storage unit separate from the database 167, and the storage unit may have a function of storing the calculation result generated by the processing unit 163a, the data input to the communication unit 161a, and the like.

[0230] [Communication units 161a and 161b] Using the communication units 161a and 161b, data can be transmitted and received between the server 151 and the terminal 152. As the communication units 161a and 161b, a hub, a router, a modem, etc. can be used. For data transmission and reception, either wired or wireless (e.g., radio waves, infrared rays, etc.) can be used.

[0231] This embodiment can be appropriately combined with other embodiments.

Explanation of Reference Numerals

[0232] S1: sentence, S2: sentence, S3: sentence, S26: sentence, STB: search text block, STD: search document, STS1: sentence, STS2: sentence, STSp: sentence, TB: text block, TB1: text block, TB2: text block, TB3: text block, TB4: text block, TB6: text block, TB7: text block, TB9: text block, TB62: text block, TD: search target document, TD1: search target document, TD2: search target document, TDn: search target document, 100: document search system, 101: input unit, 102: transmission path, 103: processing unit, 105: storage unit, 107: database, 109: output unit, 110: first target, 110(i): first target, 120: second target, 120(i): second target, 150: document search system, 151: server, 152: terminal, 161a: communication unit, 161b: communication unit, 162: transmission path, 163a: processing unit, 163b: processing unit, 165: storage unit, 167: database, 168: transmission path, 169: input / output unit

Claims

1. A document search system that searches for a specific sentence block from a plurality of sentence blocks created by splitting a plurality of documents to be searched, comprising: a processing unit; wherein the processing unit has a function of preparing a first search sentence block which is one of a plurality of search sentence blocks created by splitting a search document; uses the first search sentence block as a search condition to perform a full-text search on at least a part of the plurality of sentence blocks as a first target, and calculates a first relevance degree of each sentence block included in the first target with respect to the first search sentence block; has a function of determining a second target from among the first targets based on the height of the first relevance degree; has a function of calculating a first similarity degree between each sentence included in the first search sentence block and each sentence included in the second target; and has a function of searching for at least one sentence block similar to the first search sentence block using the first similarity degree.

2. A document search system that searches for a specific sentence block from a plurality of sentence blocks created by splitting a plurality of documents to be searched, comprising: a processing unit; wherein the processing unit has a function of preparing a first search sentence block which is one of a plurality of search sentence blocks created by splitting a search document; uses each sentence included in the first search sentence block as a search condition to perform a full-text search on at least a part of the plurality of sentence blocks as a first target, and calculates a first relevance degree of each sentence included in the first target with respect to each sentence included in the first search sentence block; has a function of determining a second target from among the sentences included in the first target based on the height of the first relevance degree; has a function of calculating a first similarity degree between each sentence included in the first search sentence block and each sentence included in the second target; and has a function of searching for at least one sentence block similar to the first search sentence block using the first similarity degree.

3. In Claim 1 or Claim 2, wherein the processing unit has a function of searching for at least one sentence block similar to the first search sentence block using a value greater than or equal to a threshold among the first similarity degrees.

Citation Information

Patent Citations

  • Method and device for retrieving similar document

    JP2004295712A

  • Computer program for retrieving relevant document and relevant document retrieving system and method

    JP2006092135A

  • Document index creating device

    JP2012104051A

  • Search device

    WO2013098886A1