Document search system

By dividing documents into blocks and calculating relevance and similarity, the method addresses the inefficiencies in identifying similar documents, improving search accuracy and efficiency, particularly for intellectual property documents.

JP7860313B2Active Publication Date: 2026-05-15SEMICON ENERGY LAB CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SEMICON ENERGY LAB CO LTD
Filing Date
2025-06-27
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing document search methods struggle to accurately identify similar documents at the block level, particularly in cases where overall similarity scores are high but partial similarities are low, leading to inefficiencies and inaccuracies in document referencing and citation processes.

Method used

A document search method that divides documents into blocks, calculates relevance and similarity at the block and sentence levels, using a first and second search text block to identify similar text blocks by performing full-text searches with relevance and similarity calculations.

Benefits of technology

This approach enables efficient and accurate identification of similar document blocks, reducing processing time and user input burden, while enhancing the precision of document searches, especially for intellectual property documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007860313000001
    Figure 0007860313000001
  • Figure 0007860313000002
    Figure 0007860313000002
  • Figure 0007860313000003
    Figure 0007860313000003
Patent Text Reader

Abstract

To provide a document search method, a document search system, a program and a storage medium for searching for similar documents with high accuracy for each block of a document.SOLUTION: A method includes: searching for a specific sentence block from among a plurality of sentence blocks formed by dividing each of a plurality of target documents; preparing a first search sentence block, which is a part of a search document; performing full-text search on at least a part of the sentence blocks, as first targets, using the first search sentence block as a search condition, to calculate first degrees of association between the first search sentence block and each of sentence blocks included in the first targets; determining, based on the first degrees of association, a second target from among the first targets; calculating first degrees of similarity between each sentence included in the first search sentence block and each sentence included in the second target; and searching for at least one sentence block similar to the first search sentence block, using the first degrees of similarity.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One aspect of the present invention relates to a document search method, a document search system, a program, and a non-temporary computer-readable storage medium. Note that one aspect of the present invention is not limited to the above technical field. Examples of the technical field of one aspect of the present invention include semiconductor devices, display devices, light-emitting devices, power storage devices, storage devices, electronic devices, lighting devices,

[0002] input devices (e.g., touch sensors, etc.), input / output devices (e.g., touch panels, etc.), their driving methods, or their manufacturing methods.

Background Art

[0003] Document search technologies for efficiently searching for target documents from a large number of documents have been actively developed. For example, Patent Document 1 discloses a similar document search method.

[0004] Similar documents may be similar as a whole to the target document, or may have extremely high similarity in some parts and extremely low similarity in other parts.

[0005] In Patent Document 1, for a target document, the detail level is calculated as an index for determining whether similar documents are similar as a whole or only partially similar.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0007] In patent application work, when creating a new specification (a later application specification), the company has previously created... The description in the previously filed specification (the specification of the prior application) may be used as a reference or reference. If a translation of the prior application's specification has already been prepared, then when preparing the translation of the later application's specification, You may refer to or cite the translation of the specification of a prior application, and you may use the translation of the specification of a later application as a reference. This can shorten the time required.

[0008] Depending on the search method for similar documents, among the documents that are calculated to have a high similarity to the target document, Even if they are not actually similar, they have a certain degree of similarity overall, so the whole document is of a certain type. Some documents may be calculated to have a high similarity score. On the other hand, the similarity of the remaining parts is extremely low. However, documents that have extremely similar parts (for example, including sentences that are exactly the same) are considered to be... The overall similarity score of the book may be calculated as low. For example, if the translated text is used as a reference, For citation purposes, the latter document is preferable to the former.

[0009] Additionally, by searching sentence by sentence, it is possible to find an exact match, but The flow of information is sometimes interrupted, and the terminology used in the specifications is not consistent. Therefore, it is desirable to be able to identify similar sections within sentence units that contain multiple sentences, such as within each chapter. .

[0010] Furthermore, when creating a new specification, one may not refer to just one specification. Therefore, In addition to specifying which specifications were used as references to create the new specification, it also specifies which part of which specification. It is desirable that by referring to the minutes, it be easy to understand which part of the new specification has been created. This is true not only for specifications but also for all documents in general. However when creating a new document, it is time-consuming and cumbersome to record in detail which parts of which documents were referenced during the process

[0011] One aspect of the present invention aims to provide a document search method that can search for similar documents for each block of a document Or, one aspect of the present invention aims to provide a document search system that can search for similar documents for each block of a document. Or One aspect of the present invention aims to provide a document search method that can search for similar documents for each block of a document with a simple input method Another aspect of the present invention aims to provide a document search method that can search for documents with high accuracy. Or, one aspect of the present invention aims to provide a document search system that can search for documents with high accuracy. Or One aspect of the present invention aims to realize document search with high accuracy, particularly for documents related to intellectual property, with a simple input method

[0012] Note that the description of these problems does not prevent the existence of other problems. One aspect of the present invention is not necessarily required to solve all of these problems. It is possible to extract other problems from the descriptions of the specification, drawings, and claims

[0013]

Means for Solving the Problems

[0014] ​​​​​​Prepare a first search text block, which is a part of the plurality of text blocks, and use the first search text block as a search condition to perform a full-text search on at least a part of the plurality of text blocks as a first target. By doing so, calculate the first relevance of each text block included in the first target to the first search text block, and based on the degree of the first relevance, determine a second target from the first target, and calculate the first similarity between each sentence included in the first search text block and each sentence included in the second target. Using the first similarity, search for at least one text block similar to the first search text block. This is a document search method.

[0015] It is preferable to create a plurality of search text blocks by splitting the search document. At this time, the first search text block is preferably one of the plurality of search text blocks.

[0016] Furthermore, prepare a second search text block, which is another part of the search document, and use the second search text block as a search condition to perform a full-text search on at least a part of the plurality of text blocks as a third target. By doing so, calculate the second relevance of each text block included in the third target to the second search text block, and based on the degree of the second relevance, determine a fourth target from the third target, and calculate the second similarity between each sentence included in the second search text block and each sentence included in the fourth target. Using the second similarity, it is preferable to search for at least one text block similar to the second search text block. At this time, the first target and the third target may be the same or different from each other.

[0017] ​​​​​​​​ Using the first similarity value above the threshold, similar text blocks to the first search text block are found. It is preferable to search for at least one lock.

[0018] One aspect of the present invention relates to a plurality of searchable documents for each of a plurality of searchable document blocks. From among the multiple text blocks created by dividing each of them, similar text blocks are selected. A document search method for searching for a document, which involves dividing the search document into multiple search documents. Create a lock and, for each of the multiple searchable text blocks, At least a portion of these will be designated as the first target, and the full text will be searched using the search text block as the search condition. By performing a search, the searchable text block is determined for each text block included in the first target. The steps involve calculating the degree of relevance and, based on the degree of relevance, selecting the second from the first target. The first step is to determine the target, and for each sentence included in the search text block, the second target is to determine the target. The steps involve calculating the similarity between each sentence and the searchable text block using the similarity. A document search method that performs the steps of: searching for at least one sentence block similar to "ku" That is the case.

[0019] One aspect of the present invention is a set of multiple documents created by dividing a set of multiple searchable documents. This is a document search method that searches for a specific text block within a block, and is used as the search document. Prepare a first searchable text block, and select at least one of several text blocks. Using the section as the first target, and using each sentence contained in the first searchable text block as a search condition, By performing a text search, the first searchable text block contains each sentence included in the first target. A first relevance score is calculated for each sentence included, and the sentences included in the first searchable text block are... Furthermore, based on the degree of relevance of the first object, the second object is determined from among the sentences included in the first object. Determined, for each sentence contained in the first search text block, and for each sentence contained in the second target A first similarity score is calculated, and the first similarity score is used to determine the similarity of the first search text block. This is a document search method that searches for at least one block of text.

[0020] It is preferable to create multiple searchable text blocks by dividing the searchable document. In this case, the first searchable text block is preferably one of several searchable text blocks. It seems so.

[0021] Furthermore, a second block of searchable text, which is another part of the searchable document, is prepared, and multiple sentences At least a portion of the block is designated as the third target and is included in the second searchable text block. By performing a full-text search using each of the sentences as search criteria, the following can be determined for each sentence included in the third target: A second relevance score is calculated for each sentence contained in the second search text block, and the second search For each sentence within a text block, based on the second level of relevance, it is included in the third target. From the sentences, a fourth target is determined, and for each sentence included in the second searchable text block, the fourth A second similarity is calculated for each sentence included in the target, and the second similarity is used to determine the second It is preferable to search for at least one text block similar to the search text block. In this case, the first object and the third object may be the same, or they may be different from each other. stomach.

[0022] Using the first similarity value above the threshold, similar text blocks to the first search text block are found. It is preferable to search for at least one lock.

[0023] One aspect of the present invention relates to a plurality of searchable documents for each of a plurality of searchable document blocks. From among the multiple text blocks created by dividing each of them, similar text blocks are selected. A document search method for searching for a document, which involves dividing the search document into multiple search documents. Create a lock and, for each of the multiple searchable text blocks, At least a portion of these will be designated as the first target, and each sentence contained within the search text block will be used as a search condition. By using this to perform a full-text search, the search text block for each sentence included in the first target is generated. The steps include calculating the relevance of each sentence contained in the block and the searchable text block. For each sentence, the second target is determined from among the sentences included in the first target, based on the degree of relevance. The first step is to search for each sentence in the search text block, and then search for the sentences in the second target. The steps involve calculating the similarity between each and using that similarity to search for similar text blocks. This is a document search method that performs the step of searching for at least one block of text.

[0024] One aspect of the present invention is a document search system having the function of performing any of the above-described document search methods. That is the case.

[0025] One aspect of the present invention is a set of multiple documents created by dividing a set of multiple searchable documents. A document search system that searches for a specific text block from within a block, wherein the processing unit The processing unit has one of several search document blocks created by dividing the search document. It has two functions: one for preparing a first searchable text block, and one for selecting a few of several text blocks. A full-text search is performed using a portion of the first target text and the first searchable text block as the search condition. By doing this, the first searchable text block for each text block included in the first target A function to calculate the first degree of relevance to K, and based on the level of the first degree of relevance, the first target A function to determine the second target from among them, and for each sentence contained in the first searchable text block, A function to calculate the first similarity to each sentence included in the second target, and using the first similarity The function searches for at least one text block similar to the first searchable text block. This is a document search system that has the following features.

[0026] One aspect of the present invention has a function that causes a processor to execute any of the above-described document search methods. This is a program. One aspect of the present invention is a non-temporary computer in which the program is stored. It is a readable memory medium.

[0027] The program is stored on a computer using various types of temporary computer-readable storage media. It may be supplied to: Temporary computer-readable storage media include electrical signals, optical signals, and electromagnetic waves. Temporary computer-readable storage media include wired wires and optical fibers. The program can be supplied to the computer via a communication channel or wireless communication channel.

[0028] One aspect of the present invention is a set of multiple documents created by dividing a set of multiple searchable documents. A program that searches for a specific text block within a block, and divides the search document The first searchable text block is one of several searchable text blocks created by dividing it. The steps involve preparing the block and selecting at least a portion of the multiple text blocks as the first target. Then, by performing a full-text search using the first searchable text block as the search condition, the first target is... The first degree of relevance of each included text block to the first searchable text block is calculated. Based on the steps taken and the degree of relevance of the first set of items, the second set of items is selected from the first set of items. The steps are to determine the second target for each sentence included in the first search text block. The steps are to calculate a first similarity to each sentence and to use the first similarity to perform a first check. The steps are to search for at least one text block similar to the searched text block, and to This is a program to be executed by the processor. In one aspect of the present invention, the program is stored It is a non-temporary computer-readable storage medium.

[0029] As non-temporary computer-readable storage media, various types of tangible storage media can be used. It is possible. As a non-temporary computer-readable storage medium, for example, RAM (Random Access Memory) Volatile memory such as DOM Access Memory, ROM (Read Only) Examples of non-volatile memory include hard disk drives. (Hard Disc Drive: HDD) and Solid State Drive (Soli d State Drive (SSD) and other recording media drives, magneto-optical disks, C Examples include D-ROMs and CD-Rs. [Effects of the Invention]

[0030] A document search method according to one aspect of the present invention, which allows searching for similar documents for each block of documents. This can be provided. According to one aspect of the present invention, similar documents can be searched for in each block of documents. A document search system can be provided. According to one aspect of the present invention, a document search system can be provided that allows for easy input of documents. For each lock, a document search method can be provided that allows searching for similar documents.

[0031] According to one aspect of the present invention, a document search method that can search for documents with high accuracy can be provided. One embodiment of this invention provides a document search system that can search for documents with high accuracy. In one embodiment, a simple input method enables highly accurate document searching, particularly for documents related to intellectual property. This can be achieved.

[0032] Furthermore, the description of these effects does not preclude the existence of other effects. One aspect of the present invention is It is not necessarily required to have all of these effects. It is possible to extract effects other than those listed above. [Brief explanation of the drawing]

[0033] [Figure 1] Figure 1 is a flowchart illustrating an example of a document search method. [Figure 2] Figure 2 shows an example of the processing stage before performing a search. [Figure 3] Figures 3A, 3B, and 3C show examples of document search methods. [Figure 4] Figures 4A, 4B, and 4C show examples of document search methods. [Figure 5] Figures 5A and 5B show examples of document search methods. [Figure 6] Figures 6A, 6B, and 6C show examples of document search methods. [Figure 7] Figures 7A, 7B, and 7C show examples of document search methods. [Figure 8] Figures 8A, 8B, and 8C show examples of document search methods. [Figure 9] Figures 9A and 9B show examples of document search methods. [Figure 10] Figure 10 is a flowchart illustrating an example of a document search method. [Figure 11] Figure 11 is a flowchart illustrating an example of a document search method. [Figure 12] Figure 12 shows an example of a document search method. [Figure 13] Figure 13 is a block view of an example of a document search system. [Figure 14] Figure 14 is a block diagram showing an example of a document search system. [Modes for carrying out the invention]

[0034] Embodiments will be described in detail with reference to the drawings. However, the present invention is not limited to the following description. Without departing from the spirit and scope of the present invention, its form and details may be modified in various ways. It will be easily understood by those skilled in the art to obtain this. Therefore, the present invention is as shown in the embodiments below. The interpretation is not limited to the content stated herein.

[0035] In the configuration of the invention described below, the same part or part having a similar function is used. The same symbol is used consistently across different drawings, and explanations of its repetition are omitted. When referring to a function, the same hatch pattern may be used, and a specific symbol may not be assigned.

[0036] Furthermore, the position, size, and extent of each component shown in the drawings are, for the sake of ease of understanding, actually The location, size, and range may not be described. Therefore, the disclosed invention must always Furthermore, it is not limited to the location, size, scope, etc., disclosed in the drawings.

[0037] (Embodiment 1) In this embodiment, a document search method according to one aspect of the present invention will be described using Figures 1 to 12. The schematic diagram of the data is just one example and is not limited to it.

[0038] One aspect of the present invention is a set of multiple documents created by dividing a set of multiple searchable documents. This is a document search method that searches for a specific text block within a larger document block.

[0039] First, prepare the first searchable text block, which is part of the document used for searching.

[0040] For example, the first searchable text block can be created by extracting a portion of the searchable document. Alternatively, the first search document block is created by splitting the search document into multiple search results. It may also be one of the search text blocks.

[0041] In one embodiment of the present invention, a document search method is used to search multiple documents, and multiple text blocks are selected in advance. Create a block of text for searching, and then, when searching, create a block of text from the search document. This allows you to search for text blocks similar to the search text block. Therefore, when using the entire document as a search criterion, or when the search target is the entire document, By comparing them, it becomes easier to understand the correspondence between similar parts.

[0042] Next, at least a portion of the multiple sentence blocks is selected as the first target, and the first search sentence is created. By performing a full-text search using blocks as search criteria, the text blocks included in the first target can be searched. Calculate the first relevance score for each of the first searchable text blocks.

[0043] The more documents to be searched, the more text blocks there will be. In one aspect of the present invention, search For each text block, you can narrow down the text block to be searched (the first target). This allows for a reduction in processing load and an increase in search speed.

[0044] Next, based on the degree of relevance of the first set of items, we select the second set of items from the first set.

[0045] In full-text search, the order of sentences and words is not considered, so the relevance score calculated will differ from the similarity score. On the other hand, sentence blocks that have words in common with the search sentence block will have a relevance value Since the relevance score increases and sentence blocks with low similarity also have low relevance scores, the similarity score should be calculated accordingly. It is possible to narrow down the target with high precision.

[0046] Next, for each sentence in the first search text block, for each sentence in the second target, Calculate the first similarity score with [the given item].

[0047] Compared to full-text search, the process of calculating similarity tends to take longer. One aspect of the present invention In this method, a second target is selected from the first target, and after narrowing down the target, the similarity is calculated. Therefore, the time required for document searches can be reduced.

[0048] Similarity can be calculated based on the degree of literal matching between sentences. Unlike full-text search, In calculating similarity, the order of words in the sentence is taken into consideration. Therefore, the first search sentence... Even if a sentence in a chapter block contains words in common with other sentences, if the word order is different, the similarity level is... The value will decrease.

[0049] Then, using the first similarity metric, we find fewer text blocks similar to the first search text block. At the very least, do one search.

[0050] As described above, by using a document search method according to one aspect of the present invention, a specific part of the document to be searched can be found. It is easy to locate similar passages in other documents.

[0051] Furthermore, in one aspect of the present invention, a document search method only requires inputting a document to be searched, and the key used for the search is Because keyword selection is unnecessary, the burden on the user is reduced, and differences in search results due to skill level are minimized. It has the advantage of being less prone to damage.

[0052] Furthermore, after narrowing down the text blocks to be searched to the first target, then the second target, By calculating similarity, the time required for document searches can be reduced.

[0053] Furthermore, full-text search uses each sentence contained in the first searchable text block as a search criterion. You may do so. In this case, for each sentence included in the first target, the first searchable text block The first relevance level is calculated for each sentence contained within the block. Then, the first searchable text block is created. For each sentence included, based on the degree of relevance of the first object, select from the sentences included in the first object. The second target will be determined.

[0054] A text block contains multiple sentences. Of the sentences contained in the text block, the first search term The majority of sentences in a text block are not necessarily similar to other sentences. Therefore, the degree of similarity To search for highly similar text blocks with high accuracy, the similarity of many text blocks is important. Calculations are required, and the time it takes to calculate the similarity can be long. To reduce the time required for output, the number of text blocks, which are the second target, should be reduced. Therefore, there is a risk of missing sentence blocks that contain sentences with a high degree of similarity.

[0055] Therefore, instead of narrowing down the text block by text, we narrow down the second target from the first target at the sentence level. This is preferable. Specifically, for each sentence included in the first searchable text block, the most relevant It is preferable to search for sentences and narrow down the target for calculating similarity at the sentence level. By narrowing down the search, compared to narrowing down the target by sentence block, it is possible to find sentences with high similarity ( The goal is to both minimize the loss of (and text blocks) and reduce the time required to calculate similarity. It is possible to measure this.

[0056] <Example of document search method 1> Figure 1 shows a flowchart of the document search method. As shown in Figure 1, one embodiment of the present invention The book search method consists of six steps, from step A1 to step A6.

[0057] Unless otherwise specified, a structure has multiple elements (documents, text blocks, or sentences). Even when explaining the composition, when explaining matters common to each element, variables and The alphabet will be omitted in the explanation. For example, search target document TD1, search target document TD2 When explaining matters common to the search target document TDn, etc., refer to the search target document TD. There are cases where this is the case.

[0058] [Pre-processing] First, we will explain the process that precedes the search using Figure 2.

[0059] In the preprocessing stage, multiple search target documents (TDs) are split to create multiple document blocks (TBs).

[0060] In the document search method of this embodiment, multiple pre-prepared documents are divided into blocks. Then, during the search, the entered search document is also divided into blocks. You can search for text blocks similar to each block.

[0061] Figure 2 shows an example of preparing n search target document TDs (where n is an integer greater than or equal to 2).

[0062] There are no particular limitations on the type of document (TD) that can be searched; various documents can be used.

[0063] Examples of searchable documents include documents related to intellectual property. Specifically, the documents include the specification, claims, and abstract used in the patent application. Examples include: Furthermore, documents related to intellectual property include patent documents (published patent gazettes, patent publications). Publications such as utility model gazettes, design gazettes, and academic papers are examples of publications issued in Japan. This applies not only to publications published in Japan, but also to publications published in countries around the world that can be used as intellectual property documents. It is possible.

[0064] In addition, the searchable document type can be books, articles, reports, columns, or other documents. Various copyrighted works, including text, may be used. Furthermore, medical documents, etc., may be used as the search target document (TD). It's okay to be there.

[0065] Furthermore, there are no particular restrictions on the language of the document; for example, Japanese, English, Chinese, Korean, etc. Which documents can be used?

[0066] The search target document TD1 shown in Figure 2 consists of x (where x is an integer greater than or equal to 2) text blocks. The block TB1(1) is divided into text blocks TB1(x)).

[0067] Furthermore, the search target document TD2 consists of y text blocks (where y is an integer greater than or equal to 2). TB2(1) is split into text block TB2(y).

[0068] Furthermore, the search target document TDn consists of z (where z is an integer greater than or equal to 2) text blocks. TBn(1) is divided into text blocks TBn(z).

[0069] For example, if the document to be searched consists of multiple chapters, by dividing it into chapters, You can also create blocks of numbers in the text.

[0070] Specifically, in the case of a patent specification, "Background, Problem, Means, and Effects," "Embodiment 1," It can be divided into "Embodiment 2," etc.

[0071] Furthermore, academic papers are often divided into sections such as "Introduction," "Research Methods," "Results," "Discussion," and "Conclusion." It is possible.

[0072] Furthermore, multiple sentence blocks may be created using all the sentences in the document to be searched. You may create multiple text blocks using only the necessary parts of the original document.

[0073] For example, if the document to be searched is a patent specification, instead of using "explanation of symbols," multiple sentence blocks... You may create a .

[0074] Preprocessing must be performed at least once before performing the document search (before performing step A1). The process may be performed multiple times depending on the application. For example, pre-processing may be performed periodically to search for data. Adding, updating, or deleting documents can improve search accuracy and usability. .

[0075] Furthermore, using multiple text blocks (TB), an index file for use in full-text search is created. It is preferable to create a log. This allows for full-text searches to be performed quickly. The structure of a DEX file is not particularly limited; for example, it could include strings, document names, and document block names. It can also contain information such as the frequency of occurrence.

[0076] Also, for example, the index file is the document TD (or document block TB) to be searched. It may also have information on whether or not a translation exists for each language. This allows, during a search, Specify conditions such as "an English translation exists" or "a Chinese translation exists." It is possible.

[0077] Next, we will explain the details of the six steps shown in Figure 1 using Figures 3 to 5.

[0078] [Step A1: Creating multiple searchable text block STBs] First, by splitting the search document STD, multiple search document blocks STB are created. (Figure 3A).

[0079] As shown in Figure 3A, the search document STD is a block of search documents containing w (where w is an integer greater than or equal to 2) The text is divided into blocks (search text block STB(1) to search text block STB(w)). It can be done.

[0080] In the document search method of this embodiment, the input search document STD is used to search multiple search documents. To separate into lock STBs, similar documents (text blocks) are sorted for each searchable text block STB. You can search for (TB).

[0081] There are no particular limitations on the searchable document (STD); various documents can be used.

[0082] Examples of searchable documents (STDs) include, for instance, documents related to intellectual property that have not yet been translated. This allows you to search for similar translated documents within the searchable TD documents, and translate them. You may refer to or quote from the text.

[0083] Additionally, the searchable document STD includes various types of documents such as books, articles, reports, columns, or texts. Copyrighted works can be used. This allows you to search for similar documents from among the searched documents (TD). You can search and check whether the searchable document STD is suspected of being plagiarized or copied. It is possible.

[0084] Additionally, medical documents can be used as searchable documents (STDs), including records of the treatment process. By using the collected medical documents to search for medical documents of similar cases, it is possible to use them as a reference for treatment. This allows us to consider what course the patient's condition might take in the future.

[0085] [Step A2: Select the search text block STB(i)] Next, the search text block STB that performs the search is selected from among the w search text block STBs. (i) Select an integer between 1 and w (i is greater than or equal to 1).

[0086] Note that if you want to perform a search on only one searchable text block STB, proceed to step A1. By extracting the necessary parts from the search document STD, the search text block ST You may create B.

[0087] Furthermore, when performing a search on multiple searchable text blocks (STBs), you must search each one individually. You can perform a subsequent search (see Example 3 of Document Search Methods), or you can perform multiple searches in parallel (document search). (See Example 4 of the search method) and a combination of sequential and parallel processing may be used for searching.

[0088] In the document search method of this embodiment, for each searchable document block STB, similar document blocks are found. Because you can search for TB, you can search for similar parts of the search document STD. The location of the relevant section in the document TD can be identified accurately and easily.

[0089] [Step A3: Calculating the relevance of the search text block STB(i)] Next, we calculate the relevance to the search text block STB(i).

[0090] Specifically, by using the searchable text block STB(i) as the search condition to perform a full-text search For each of the text blocks TB to be searched, the search text block STB(i) Calculate the degree of relevance.

[0091] Here, for all text blocks TB, the relationship with the search text block STB(i) The degree of repetition may be calculated, and for some text blocks TB, the search text block STB( You may also calculate the degree of relevance to i).

[0092] For example, in the case of a patent specification, I searched for similar documents regarding "background, problem, means, and effects." In that case, you only need to search the "background, issues, means, and effects" of the documents you are searching. "Embodiment 1," etc., can be excluded from the search.

[0093] Furthermore, if you want to find similar documents regarding "Embodiment 1," the form of each implementation in the searched documents will be... You can search for "state" and exclude "background, issues, means, and effects." Furthermore, if you want to find similar documents for which "an English translation exists", search for "documents for which an English translation exists". Each embodiment of the document containing the "existing" information can be included in the search.

[0094] In full-text search, the text block TB used to calculate relevance is, for example, an index file. The selection is made automatically based on the information contained in the document. Alternatively, enter the search document STD. You may also specify a text block TB for which relevance is to be calculated.

[0095] In this way, the text block to be searched is selected according to the search text block STB(i). By making these changes, the amount of processing required can be reduced, and the time it takes to search for documents can be shortened.

[0096] In example 1 of the document search method, the search text block STB(i) is used as one of the search criteria for full-text search. This shows how it will be used as a case. Furthermore, as will be described later, the search text block STB(i) contains Each included sentence may be used as a search criterion for a full-text search (see Example 2 of the document search method). In other words, the number of search conditions is equal to the number of sentences contained in the search text block STB(i). That's good too.

[0097] There are no particular restrictions on the full-text search method; sequential search, index search, etc., can be used.

[0098] In particular, index search can search even when the number of document blocks (TB) to be searched is large. This is preferable because it does not significantly reduce speed.

[0099] In index search, the text blocks TB to be searched are scanned in advance, Prepare an index file that enables fast searching.

[0100] There are no particular restrictions on the method of extracting the strings that make up the index file, including word segmentation. (Separating words with spaces), morphological analysis, N-gram (N-character indexing method, N The gram system (also known as the gram system) can be used.

[0101] In particular, N-gram is advantageous for exact match searches compared to morphological analysis, and for specialized terminology. This is preferable because it minimizes the likelihood of problems arising from new words and abbreviations.

[0102] For example, the correlation can be calculated using TF-IDF (Term Frequency-Inverse). It is preferable to use (Document Frequency). The TF value is The IDF value represents the frequency of occurrence of each word within a given block of text. This indicates the degree to which a word appears in a concentrated area. The more often a word appears in a single sentence block, the more concentrated it is. The TF value of the word in question will be high in that sentence block. It appears in many sentence blocks. Words with a low IDF value tend to have a higher IDF value if they appear only in certain sentence blocks. By calculating the product of the TF value and IDF value of each word, we can determine how that word characterizes a block of text. It is possible to calculate a score to determine whether something is a word or not.

[0103] Furthermore, the calculation of the degree of relevance is not limited to the method using TF-IDF.

[0104] For example, Apache Lucene, an open-source search engine library This allows you to perform a full-text search.

[0105] Figure 3B shows an example of calculating the relevance to the searchable text block STB(1). The first target 110(1), which is the search target, is the first sentence that each search target document TD has An example of block TB(1) is shown.

[0106] [Step A4: Determine the second target 120(i) from the first target 110(i)] Next, based on the degree of relevance, select the second object 120(i) from the first object 110(i). ) will be decided.

[0107] The number of text blocks TB included in the second target 120(i) is not particularly limited. The second target 120(i) becomes the target for calculating similarity in the next step. Compared with full-text search, the process of calculating similarity tends to take a long time. By determining the second target 120(i) from the first targets 110(i) and calculating the similarity after narrowing down the targets, the time required for document search can be shortened.

[0108] For example, by sorting the results of the full-text search in step A3 in descending order of relevance, it is possible to identify the text blocks TB with high relevance to the search text block STB(i). can be done.

[0109] In FIG. 3C, an example of using the top 10 text blocks TB with high relevance to the search text block STB(1) as the second target 120(1) is shown. In FIG. 3C, as an example, the text block TB4(1) is ranked 1st (Rank 1), the text block TB1(1) is ranked 2nd ( Rank 2), and the text block TB9(1) is ranked 10th (Rank 10) is shown. case.

[0110] [Step A5: Calculation of similarity for the search text block STB(i)] Next, the similarity for the search text block STB(i) is calculated. Specifically, for each sentence included in the search text block STB(i), the similarity with each sentence included in the second target 120(i) is calculated.

[0111] ​​​​​ For example, using the diff algorithm, which finds the difference between documents, you can calculate similarity. It is possible.

[0113] First, as shown in Figure 4A, the first sentence STS1 of the search text block STB(1) and The similarity to each sentence included in the second target 120(1) is calculated.

[0114] Next, as shown in Figure 4B, the second sentence STS2 of the search text block STB(1) and The similarity to each sentence included in the second target 120(1) is calculated. Similarly, the search sentence The relationship between each sentence in chapter block STB(1) and each sentence in the second object 120(1) Calculate the similarity score.

[0115] Then, as shown in Figure 4C, the last sentence STSp(p The similarity is calculated up to an integer greater than or equal to 1, and the search text block STB(1) contains For all sentences, calculate the similarity to each sentence included in the second object 120(1). To output. Note that Figure 4C shows an example where p is an integer greater than or equal to 3.

[0116] Furthermore, the calculation of similarity between multiple sentences in the search text block STB(1) is performed in parallel. It is also possible. For example, the process shown in Figure 4A, the process shown in Figure 4B, and the process shown in Figure 4C are all They may be performed in parallel.

[0117] By using the calculated similarity score, similar text blocks to the search text block STB(1) can be found. We can find the value of TB.

[0118] For example, in each text block TB, for each sentence of the search text block STB(1) The sum of the similarity scores of the sentences with the highest similarity is calculated, and this sum is used in the search text block STB(1). By dividing by the number of sentences, the searchable text block STB(1) of the text block TB is determined. The normalized similarity can be calculated.

[0119] In Figure 5A, in text block TB4(1), 1 of the search text block STB(1) The sentence with the highest similarity to the second sentence STS1 is the first sentence S1 (similarity is 1). The sentence with the highest similarity to the second sentence STS2 is the second sentence S2 (similarity: 0.9). Yes, the sentence with the highest similarity to the last sentence STSp is the third sentence S3 (similarity 0.5). ) These p similarities are added together and divided by the number of sentences p, resulting in the sentence block TB4(1 The normalized similarity of the searchable text block STB(1) can be calculated.

[0120] Furthermore, using similarity scores above a certain threshold between sentences can improve search accuracy. Therefore, it is preferable. For example, if the threshold is 0.8, the text block TB4 shown in Figure 5A (1) The sentence S3, which has the highest similarity to the last sentence STSp, has a similarity of 0.5. Therefore, it is not used (it is treated as 0) when calculating the sum of similarities.

[0121] [Step A6: Outputting Results] Furthermore, the text block TB has a high normalization similarity to the search text block STB(i). Outputs.

[0122] Figure 5B shows an example of text blocks (TB) arranged in order of normalization similarity. Furthermore, an example of expressing the normalized similarity as a percentage is shown as the score.

[0123] In the full-text search performed in Step A3, the order of sentences and words is not considered, so the calculated relation The degree of similarity is different from the degree of kinship. By calculating the degree of similarity in step A5, step A4 (Figure 3) The 10 text blocks TB determined as the second target 120(1) in C) are used as search texts. The blocks can be arranged in order of their similarity to block STB(1) (Figure 5B).

[0124] As described above, the search document STD is divided into search text blocks STB, and similar text blocks are selected. By searching for locks, similar documents (text blocks) are found for the search text block STB. You can search for (TB). This allows you to use the entire search document STD as a search criterion. Compared to cases where the search target is the entire document, it is possible to understand the correspondence between similar parts. This makes it easier.

[0125] Furthermore, after narrowing down the text blocks to be searched to the first target, then the second target, By calculating similarity, the time required for document searches can be reduced.

[0126] <Example of document search method 2> Next, we will explain the modified examples from step A3 onwards using Figures 6 to 9. Specifically, for search purposes When using each sentence contained in the text block STB(i) as a search criterion for full-text search, I will explain.

[0127] [Step A3: Calculating the relevance of the search text block STB(i)] In step A3 of example 2 of the document search method, the search text block STB(i) contains A full-text search is performed using each sentence as a search criterion. This allows each sentence included in the search target to be found. Then, the relevance of each sentence in the searchable text block STB(i) is calculated.

[0128] Here, for all text blocks TB, the search text block STB(i) contains The relevance of each sentence may be calculated, and for some sentence blocks TB, the search sentence block The degree of relevance for each sentence included in lock STB(i) may also be calculated.

[0129] By changing the text block to be searched according to the search text block STB(i) This reduces the amount of processing required and shortens the time it takes to search for documents.

[0130] The full-text search method and the method for calculating relevance should be the same as in Example 1 of the document search method. It is possible.

[0131] First, as shown in Figure 6A, the first sentence STS1 of the search text block STB(1) is searched. By using this as a search condition and performing a full-text search, one of the sentences included in the first target 110(1) The degree of relevance to the sentence STS1 is calculated. Note that the sentence included in the first target 110(1) This refers to sentences that make up the multiple sentence block TB included in the first object 110(1).

[0132] Next, as shown in Figure 6B, the second sentence STS2 of the search text block STB(1) is searched. By using the search conditions to perform a full-text search, the two sentences included in the first target 110(1) The relevance of the text STS2 to the eye is calculated. Similarly, the relevance of the search text block STB(1) Calculate the degree of relevance for each sentence.

[0133] Then, as shown in Figure 6C, the last sentence STSp(p By calculating the degree of relevance up to an integer of 2 or more, the sentences included in the first target 110(1) The relevance of each sentence in the searchable text block STB(1) is calculated. Figure 6C shows an example where p is an integer greater than or equal to 3.

[0134] Furthermore, full-text searches using each sentence in the search text block STB(1) as search criteria are performed in parallel. It is also acceptable. For example, the process shown in Figure 6A, the process shown in Figure 6B, and the process shown in Figure 6C are, All processes may be carried out in parallel.

[0135] [Step A4: Determine the second target 120(i) from the first target 110(i)] Next, for each sentence contained in the search text block STB(i), based on the degree of relevance, The second object 120(i) is determined from the sentences included in the first object 110(i).

[0136] The number of sentences included in the second object 120(i) is not particularly limited. ) will be used to calculate similarity in the next step. Compared to full-text search, calculating similarity is This process tends to take a long time. From the first target 110(i), the second target 12 By determining 0(i) and narrowing down the target, the similarity is calculated, which determines when a document search is performed. The time can be shortened.

[0137] For example, by sorting the full-text search results in step A3 in order of relevance, Identify sentences that are highly relevant to each sentence contained in the searchable text block STB(i). It is possible.

[0138] Figure 7A shows the high relevance of the first sentence STS1 in the search text block STB(1). An example is shown where the top 300 sentences are used as the second target 120(1)(STS1). Figure 7 In case A, for example, the first sentence TB4(1)_S1 of sentence block TB4(1) is ranked 1st. (Rank 1), the first sentence TB3(1)_S1 of sentence block TB3(1) is ranked 2nd ( Rank 2), and the sixth sentence TB6(1)_S6 of sentence block TB6(1) is This shows the case where the rank is 300th.

[0139] Figure 7B shows the high relevance of the second sentence STS2 in the search text block STB(1). An example is shown where the top 300 sentences are used as the second target 120(1)(STS2). Figure 7 In section B, for example, the second sentence TB1(1)_S2 of sentence block TB1(1) is ranked 1st. (Rank 1), the second sentence TB3(1)_S2 of sentence block TB3(1) is ranked 2nd ( Rank 2), and the eighth sentence of sentence block TB62(1)_S This shows the case where 8 is ranked 300th.

[0140] Then, as shown in Figure 7C, the last sentence STSp of the search text block STB(1) The second set of 120(1)(STSp) was selected as the top 300 sentences with the highest relevance. In Figure 7C, as an example, the ninth sentence of text block TB2(1)_ S9 is ranked 1st (Rank 1), the 8th sentence of sentence block TB6(1) TB6(1)_S 8 is in 2nd place (Rank 2), and the 12th sentence of sentence block TB7(1) TB7( 1) Shows the case where _S12 is ranked 300th (Rank 300). As above, for search purposes For all sentences contained in text block STB(1), the second target 120( Determine 1). Similarly, for all sentences contained in the search text block STB(i) Based on their degree of relevance, the following sentences were selected from among the sentences included in the first object 110(i): Determine the target 120(i) for 2.

[0141] [Step A5: Calculating similarity to the search text block STB(i)] Next, the similarity to the search text block STB(i) is calculated. Specifically, the search For each sentence contained in the text block STB(i), the sentence contained in the second target 120(i) Calculate the similarity between each item.

[0142] The method for calculating similarity can be the same as in Example 1 of the document search method.

[0143] First, as shown in Figure 8A, the first sentence STS1 of the search text block STB(1) and The similarity to each sentence in the second target 120(1)(STS1) is calculated.

[0144] Next, as shown in Figure 8B, the second sentence STS2 of the search text block STB(1) and The similarity to each sentence in the second target 120(1)(STS2) is calculated. This includes each sentence in the search text block STB(1) and the sentences included in the second target 120(1). Calculate the similarity between each item.

[0145] Then, as shown in Figure 8C, up to the last sentence STSp of the search text block STB(1) By calculating the similarity, all sentences contained in the search text block STB(1) are compared. Then, the similarity to each sentence included in the second target 120(1) is calculated.

[0146] Furthermore, the calculation of similarity between multiple sentences in the search text block STB(1) is performed in parallel. It is also possible. For example, the process shown in Figure 8A, the process shown in Figure 8B, and the process shown in Figure 8C are all They may be performed in parallel.

[0147] By using the calculated similarity score, similar text blocks to the search text block STB(1) can be found. We can find the value of TB.

[0148] For example, in each text block TB, for each sentence of the search text block STB(1) The sum of the similarity scores of the sentences with the highest similarity is calculated, and this sum is used in the search text block STB(1). By dividing by the number of sentences, the searchable text block STB(1) of the text block TB is determined. The normalized similarity can be calculated.

[0149] In Figure 9A, in text block TB4(1), 1 of the search text block STB(1) The sentence with the highest similarity to the second sentence STS1 is the first sentence S1 (similarity is 1). The sentence with the highest similarity to the second sentence STS2 is the second sentence S2 (similarity: 0.90). Therefore, by adding up the highest similarity scores for each of the p sentences and dividing by the number of sentences p, , the normalized similarity of text block TB4(1) to search text block STB(1) This can be calculated. Note that in text block TB4(1), the 26th sentence S26 Also, the similarity to the first sentence STS1 of the search text block STB(1) is high (similar Since the similarity score (0.80) is lower than that of the first sentence S1, the similarity value of S26 is not used.

[0150] Furthermore, using similarity scores above a certain threshold between sentences can improve search accuracy. Therefore, it is preferable. In the text block TB9(1) shown in Figure 9A, the search text block The sentence with the highest similarity to the first sentence STS1 of the query sentence block STB(1) is the second sentence S2 (similarity is 0.70), and the sentence with the highest similarity to the second sentence STS2 is the first sentence S1 (similarity is 0.60), and the sentence with the highest similarity to the last sentence STSp is the third sentence S3 (similarity is 0.60). When not using a threshold value, the similarity values of these three sentences are used to calculate the sum of the highest similarities for each of the p sentences. On the other hand, for example, when the threshold value is 0.8, since the similarity values of these three sentences are less than the threshold value, they will not be used (considered as 0) when calculating the sum of similarities.

[0151] [Step A6: Output of Results] Then, output the sentence block TB with a high normalized similarity to the query sentence block STB(i).

[0152] Figure 9B is an example of arranging sentence blocks TB in descending order of normalized similarity. Also, as an example of Score, an example of expressing the normalized similarity as a percentage is shown.

[0153] In Example 2 of the document search method, for each sentence included in the query sentence block STB(i), a sentence that becomes the second target 120(i) is determined from the first target 110(i). Therefore, among the sentences included in the sentence block TB, only the sentences with a high relevance to the sentences included in the query sentence block STB(i) can calculate the similarity with the sentences included in the query sentence block STB(i). By narrowing down the target at the sentence level, it is possible to suppress the omission of sentences (and sentence blocks) with high similarity compared to the case of narrowing down the target at the sentence block level, and it is possible to shorten the time required to calculate the similarity. Also, in fact, sentences that are not actually similar sentence blocks ​​​​​​​​​​This prevents the similarity of locked TBs from becoming too high.

[0154] For example, by using document search method example 2, in document search method example 1 (Figure 5B), the top 1 The sentence blocks TB7(1), TB3(1), and TB6(1) that did not rank 0th are among the top 10. It is also possible that they could end up in a lower position (Figure 9B).

[0155] Example 2 of the document search method has significantly lower similarity in the remaining parts compared to Example 1 of the document search method. However, sentence blocks that have extremely high similarity (for example, sentences that are an exact match) This allows for a higher calculation of similarity.

[0156] <Example 3 of document search methods> Next, the system searches for similar text blocks sequentially across multiple searchable text blocks (STBs). The method will be explained. Note that in example 3 of the document search method, all searchable text blocks ST are used. Regarding B, an example of searching for similar sentence blocks is shown, but it is not limited to this, and some searches may be performed. For the searchable text block STB, you may also search for similar text blocks. Figure 10 shows: A flowchart of the document search method is shown.

[0157] Note that the processing steps prior to performing the search are the same as in Example 1 of the document search method, so please do not explain further. Omit it.

[0158] [Step B1: Creating multiple searchable text blocks STB(1) to STB(w)] First, by splitting the search document STD, multiple search document blocks STB are created. Here, we have w search text blocks (where w is an integer greater than or equal to 2) (search text block ST). This example shows how to split B(1) into a searchable text block STB(w)). Step B1 is: It can be performed in the same manner as step A1 shown in FIG. 3A.

[0159] [Step B2: Selection of the search text block STB(i) (i = 1)] Next, from among the w search text blocks STB, a search text block STB (i) (where i is an integer from 1 to w) is selected.

[0160] Note that for some or all of the search text blocks STB, the order of searching for similar text blocks is not particularly limited.

[0161] In Example 3 of the document search method, an example of performing the search in order from the search text block STB(1) is shown. Therefore, in step B2, i = 1 is selected.

[0162] [Step B3: Calculation of the relevance for the search text block STB(i)] Next, the relevance for the search text block STB(i) is calculated.

[0163] Since i = 1 is selected in step B2, in the first step B3, the relevance for the search text block STB(1) is calculated. The first step B3 can be performed in the same manner as step A3 shown in FIG. 3B.

[0164] [Step B4: Determination of the second target 120(i) from among the first targets 110(i)] Next, based on the degree of relevance, the second target 120(i ) is determined from among the first targets 110(i).

[0165] Since i = 1 is selected in step B2, in the first step B4, based on the degree of relevance, the second target 120(1) is determined from among the first targets 110(1). The first step ​​Step B4 can be performed in the same way as step A4 shown in Figure 3C.

[0166] [Step B5: Calculating similarity to the search text block STB(i)] Next, the similarity to the search text block STB(i) is calculated. Specifically, the search For each sentence contained in the text block STB(i), the sentence contained in the second target 120(i) Calculate the similarity between each item.

[0167] Because i=1 was selected in step B2, in the first step B5, the search text block was generated. The similarity to STB(1) is calculated. The first step B5 is as shown in Figures 4A to 4C and This can be done in the same way as step A5 shown in Figure 5A.

[0168] [Step B6: Have similarity scores been calculated for all search text blocks (STB) (i=w ?)] The above process from step B3 to step B5 is applied to all searchable text blocks (STB). Then proceed in order. If there is a search text block STB for which similarity has not been calculated, then Return to step B3 via step B7. Then, apply to all searchable text blocks STB. If the similarity is calculated, proceed to step B8.

[0169] [Step B7: Add 1 to i (i = i + 1)] When returning from step B6 to step B3, step B7 is created by adding 1 to i. Then, the second step B3-B5 is performed on the search text block STB(2). As shown above, step B is performed until the similarity is calculated for the search text block STB(w). Repeat steps 3-B5.

[0170] [Step B8: Outputting Results] Then, it outputs text blocks TB with a high standardization similarity to each search text block STB. To exert force.

[0171] Figure 12 shows the text blocks TB sorted by search text block STB in descending order of normalization similarity. This is an example of how they are arranged. Furthermore, as shown in Figure 5B, a value indicating the degree of similarity is generated. You may use force.

[0172] As described above, similar text blocks were searched sequentially for each searchable text block STB. Furthermore, by outputting all results, each search document block STB in the search document STD is selected. This allows you to search for similar documents (text blocks TB).

[0173] <Example 4 of document search methods> Next, we search for similar text blocks in parallel across multiple searchable text block STBs. This section explains how to do this. Note that in example 4 of the document search method, all searchable text blocks are used. The following example shows how to search for similar text blocks using STB, but it is not limited to this example. For the searchable text block STB, you may also search for similar text blocks. Figure 11 The flowchart below shows the document search method.

[0174] Note that the processing steps prior to performing the search are the same as in Example 1 of the document search method, so please do not explain further. Omit it.

[0175] [Step C1: Creating multiple searchable text block STBs] First, by splitting the search document STD, multiple search document blocks STB are created. Here, we have w search text blocks (where w is an integer greater than or equal to 2) (search text block ST). This example shows how to split B(1) into a searchable text block STB(w)). Step C1 is: This can be done in the same way as step A1 shown in Figure 3A.

[0176] The subsequent steps C2-C5 process applies to two or more searchable text blocks (STBs) in parallel. It can be done in columns. In example 4 of the text search method, there are w search text blocks STB. Next, I will show an example of performing the process in parallel.

[0177] [Step C2(i): Select the search text block STB(i)] Next, the search text block STB that performs the search is selected from among the w search text block STBs. (i) Select an integer between 1 and w (i is greater than or equal to 1).

[0178] In step C2(1) shown in Figure 11, select i=1. In parallel with step C2(1) In step C2(2), which is performed, i=2 is selected, and in step C2(w), i=w Select this option.

[0179] [Step C3(i): Calculation of relevance for the search text block STB(i)] Next, we calculate the relevance to the search text block STB(i).

[0180] In step C3(1) shown in Figure 11, the relevance to the search text block STB(1) is Calculate the value. Step C3(1) can be performed in the same way as step A3 shown in Figure 3B. ru.

[0181] Step C3(2), which is performed in parallel with step C3(1), involves a searchable text block S. The relevance to TB(2) is calculated, and in step C3(w), the search text block ST is used. Calculate the degree of relevance to B(w).

[0182] [Step C4(i): Determine the second target 120(i) from the first target 110(i) ] Next, based on the degree of relevance, select the second object 120(i) from the first object 110(i). ) will be decided.

[0183] In step C4(1) shown in Figure 11, based on the degree of relevance, the first target 110(1 From among them, the second target 120(1) is determined. Step C4(1) is as shown in Figure 3C. It can be done in the same way as Step A4.

[0184] Step C4(2), which is performed in parallel with Step C4(1), is based on the degree of relevance. Then, the second target 120(2) is determined from the first target 110(2), and step C4( w) Based on the degree of relevance, the second set of 120 items will be selected from the first set of 110 items (w). Determine (w).

[0185] [Step C5: Calculation of similarity for the search text block STB(i)] Next, the similarity to the search text block STB(i) is calculated. Specifically, the search For each sentence contained in the text block STB(i), the sentence contained in the second target 120(i) Calculate the similarity between each item.

[0186] In step C5(1) shown in Figure 11, the similarity to the search text block STB(1) is calculated. Calculate the value. Step C5(1) is the same as step A5 shown in Figures 4A-4C and 5A. It can be done in this way.

[0187] Step C5(2), which is performed in parallel with step C5(1), involves a searchable text block S. The similarity to TB(2) is calculated, and in step C4(w), the search text block ST Calculate the similarity to B(w).

[0188] [Step C6: Output of Results] Then, it outputs text blocks TB with a high standardization similarity to each search text block STB. To exert force.

[0189] Figure 12 shows the text blocks TB sorted by search text block STB in descending order of normalization similarity. This is an example of the arrangement. Note that, as shown in Figure 5B (Score), a value indicating the degree of similarity is output. You may do so.

[0190] As described above, after searching for text blocks similar to each searchable text block STB in parallel, By outputting all results, for each search text block STB in the search document STD This allows you to search for similar documents (text blocks TB).

[0191] As described above, in the document search method of this embodiment, similar document blocks to the search document block are searched for. By searching for locks, you can find similar sections in the search document. This allows for highly accurate searches. This is possible when using the entire document as a search criterion. Furthermore, it is easier to understand the correspondence between similar sections compared to when the entire document is being searched. This is the result.

[0192] Furthermore, in the document search method of this embodiment, full-text search results are used to search text blocks. The target of the similarity calculation is narrowed down. This reduces the time spent on document searches. It is possible.

[0193] This embodiment can be appropriately combined with other embodiments. Furthermore, this specification Furthermore, if multiple configuration examples are shown within a single embodiment, the configuration examples may be combined as appropriate. It is possible to do so.

[0194] (Embodiment 2) In this embodiment, Figures 13 and 14 are used to describe a document search system according to one aspect of the present invention. I will explain.

[0195] The document search system of this embodiment uses the document search method shown in Embodiment 1 to search for documents It can be searched. Specifically, it searches pre-prepared blocks of text. The system searches for documents (text blocks) similar to the entered search document (search text block). It can be searched.

[0196] <Example of document search system configuration 1> Figure 13 shows a block diagram of the document search system 100. Note that the drawings attached to this specification... So, let's classify the components by function and show them as independent blocks in a block diagram. However, it is difficult to completely separate the actual components by function, and one component is It may involve multiple functions. Also, one function may involve multiple components. For example, the processing performed in processing unit 103 may be executed on different servers depending on the process. It can happen.

[0197] The document search system 100 has at least a processing unit 103. (Figure 13 shows the document search) System 100 further includes an input unit 101, a transmission line 102, a storage unit 105, and a database. It has a 107 and an output section 109.

[0198] [Input section 101] The input unit 101 is supplied with searchable documents STD from outside the document search system 100. The search document STD supplied to the input unit 101 is transmitted via the transmission line 102 to the processing unit 103. It is supplied to the storage unit 105 or the database 107.

[0199] [Transmission path 102] The transmission line 102 has the function of transmitting various data. Input unit 101, processing unit 103, Data transmission and reception between the memory unit 105, the database 107, and the output unit 109 is performed via transmission path 1 This can be done via 02. For example, search document STD, search text block STB Data such as the search target document TD and the document block TB are transmitted via the transmission path 102. It will be received.

[0200] [Processing step 103] The processing unit 103 receives data supplied from the input unit 101, storage unit 105, database 107, etc. It has the function of performing calculations using data. The processing unit 103 stores the calculation results in the storage unit 105 This can be supplied to the database 107, output unit 109, etc.

[0201] The processing unit 103 uses a transistor having a metal oxide in the channel formation region. Preferred. Because the transistor has an extremely low off-current, the transistor can be used as a memory element. It is used as a switch to hold the charge (data) that flows into a capacitive element that functions as such. This ensures that data can be retained for a long period of time. By using it in at least one of the registers and cache memory of 103, The processing unit 103 is operated only when necessary, and in other cases, the information of the previous processing is stored in the memory element. By putting it into standby mode, the processing unit 103 can be turned off. In other words, normally This enables power-saving computing, allowing for lower power consumption in document search systems. ru.

[0202] In this specification, etc., when an oxide semiconductor or metal oxide is used in the channel formation region. Transistors are called Oxide Semiconductor transistors, or OS transistors. It is called a transistor. The channel formation region of an OS transistor may contain a metal oxide. preferable.

[0203] In this specification, metal oxide refers to metals in a broad sense. It is an oxide. Metal oxides are oxide insulators and oxide conductors (including transparent oxide conductors). Oxide semiconductors (also called OS) They are classified into the following categories. For example, when a metal oxide is used in the semiconductor layer of a transistor, the metal Oxides are sometimes referred to as oxide semiconductors. In other words, metal oxides have amplification and rectification effects. , and if it has at least one switching action, the metal oxide is a metal oxide A semiconductor (metal oxide semiconductor), abbreviated as OS It is possible.

[0204] The metal oxide in the channel-forming region preferably contains indium (In). If the metal oxide in the Nell-forming region is an indium-containing metal oxide, the OS Transis The carrier mobility (electron mobility) of the ion becomes higher. Also, the metallic acid present in the channel-forming region The oxide is preferably an oxide semiconductor containing element M. Element M is aluminum (Al). It is preferable that it be gallium (Ga) or tin (Sn). Other applicable elements M The elements include boron (B), silicon (Si), titanium (Ti), iron (Fe), and nickel. Kel (Ni), Germanium (Ge), Yttrium (Y), Zirconium (Zr), Mo Ribdenum (Mo), Lanthanum (La), Cerium (Ce), Neodymium (Nd), Hafniu Examples include fluorine (Hf), tantalum (Ta), and tungsten (W). However, as element M... In some cases, it is acceptable to combine multiple of the aforementioned elements. Element M, for example, can be combined with oxygen. It is an element with high bonding energy. For example, its bonding energy with oxygen is higher than that of indium. It is an element. Furthermore, the metal oxides that the channel-forming region contains include zinc (Zn). This is preferable. Zinc-containing metal oxides may be prone to crystallization.

[0205] The metal oxides present in the channel-forming regions are not limited to indium-containing metal oxides. The semiconductor layer is made of materials such as zinc tin oxide and gallium tin oxide, which do not contain indium. These included metal oxides containing zinc, metal oxides containing gallium, and metal oxides containing tin. That's fine.

[0206] Furthermore, the processing unit 103 may use a transistor that includes silicon in its channel formation region. stomach.

[0207] Furthermore, the processing unit 103 includes a transistor containing an oxide semiconductor in the channel formation region, and a channel It is preferable to use a transistor containing silicon in the Nel-forming region in combination with the other transistor.

[0208] The processing unit 103 is, for example, an arithmetic circuit or a central processing unit (CPU). It has an operating unit, etc.

[0209] The processing unit 103 includes a DSP (Digital Signal Processor) and a GP (Ground Processing Unit). It has a microprocessor such as U (Graphics Processing Unit). It is acceptable to do so. Microprocessors are FPGAs (Field Programmable Arrays). Field Programmable Array), FPAA (Field Programmable A Programmable Logic Dev (PLD) such as a rectangular array The configuration may be implemented by ice. The processing unit 103 is determined by the processor. By interpreting and executing instructions from various programs, various data processing and programmatic processes are performed. It is possible to perform the action. The programs that can be executed by the processor are those that the processor possesses. It is stored in at least one of the memory area and the storage unit 105.

[0210] The processing unit 103 may have main memory. The main memory may be volatile memory such as RAM. It has at least one of Mori and non-volatile memory such as ROM.

[0211] For example, RAM can be DRAM (Dynamic Random Access Memory). mory), SRAM (Static Random Access Memory), etc. This is used, and a memory space is virtually allocated and used as the workspace for the processing unit 103. The operating system and application programs stored in the memory unit 105. Program modules, program data, and lookup tables are used for execution. These are then loaded into RAM. These data, programs, and programs loaded into RAM are then loaded into RAM. Each program module is directly accessed and operated by the processing unit 103.

[0212] The ROM contains BIOS (Basic Input / Output) which does not require rewriting. It can store the System and firmware, etc. As for ROM, SCRROM, OTPROM (One Time Programmable Read) Only Memory), EPROM (Erasable Programmable Examples include Read Only Memory. EPROMs include ultraviolet light UV-EPROM (Ultra-Violet) enables the erasure of stored data by irradiation. Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmability) Examples include Read Only Memory (e) and flash memory.

[0213] [Storage section 105] The memory unit 105 has the function of storing the program to be executed by the processing unit 103. The memory unit 105 stores the calculation results generated by the processing unit 103 and the data input to the input unit 101. It may also have a function to store things like "ta".

[0214] The storage unit 105 has at least one of volatile memory and non-volatile memory. The unit 105 may have, for example, volatile memory such as DRAM or SRAM. Part 105 is, for example, ReRAM (Resistive Random Access). Memory (also called resistive random-access memory), PRAM (Phase change R) andom Access Memory), FeRAM (Ferroelectric Random Access Memory), MRAM (Magnetoresis (also known as magnetically resistive random access memory) Alternatively, it may have non-volatile memory such as flash memory. Also, storage unit 105 These include hard disk drives (HDDs) and solid-state drives. Recording media such as state drives (Solid State Drive: SSD) It's fine if they have live performances.

[0215] [Database 107] Database 107 contains at least the data of the searchable document TD and document block TB. It has the function of storing data. In addition, the database 107 stores calculations generated by the processing unit 103. The system may also have a function to store the results and data entered into the input unit 101. Note that the storage unit 105 and the database 107 do not necessarily have to be separated from each other. The document search system has the functions of both a storage unit 105 and a database 107. It may have units.

[0216] Furthermore, the memory of the processing unit 103, the storage unit 105, and the database 107 is Therefore, it can be considered an example of a non-temporary computer-readable storage medium.

[0217] [Output section 109] The output unit 109 has the function of supplying data to the outside of the document search system 100. If so, the calculation results from the processing unit 103 can be supplied to an external source.

[0218] <Example of document search system configuration 2> Figure 14 shows a block diagram of the document search system 150. The document search system 150 is a... It has a server 151 and a terminal 152 (such as a personal computer).

[0219] Server 151 includes a communication unit 161a, a transmission line 162, a processing unit 163a, and Database 1 It has 67. Although not shown in Figure 14, the server 151 also has a storage unit, an input / output unit, and It may have any of the following.

[0220] Terminal 152 includes a communication unit 161b, a transmission line 168, a processing unit 163b, a storage unit 165, and input It has an output unit 169. Although not shown in Figure 14, terminal 152 further has a database They may have, etc.

[0221] Users of the document search system 150 can access the search document STD from terminal 152 to server 15 Enter into 1. The search document STD is sent from communication unit 161b to communication unit 161a.

[0222] The search document STD received by the communication unit 161a is transmitted via the transmission line 162 to database 1 67 or stored in the memory unit (not shown). Alternatively, the search document STD is stored in the communication unit 1 The power may be supplied directly from 61a to the processing unit 163a.

[0223] The creation of a searchable text block STB, calculation of relevance, and similarity as described in Embodiment 1. Each of these calculations requires high processing power. The processing unit 163a of server 151 This has higher processing power than the processing unit 163b of terminal 152. Therefore, these The processing is preferably carried out in the processing unit 163a.

[0224] Then, the processing unit 163a generates the search results. The search results are transmitted via the transmission line 162. The results are then stored in database 167 or a storage unit (not shown). Alternatively, the search results are Alternatively, the processing unit 163a may supply the data directly to the communication unit 161a. After that, the server 15 From 1, the search results are output to terminal 152. The search results are sent from communication unit 161a to communication unit It will be sent to 161b.

[0225] [Input / output section 169] Data is supplied to the input / output unit 169 from outside the document search system 150. 169 has the function of supplying data to the outside of the document search system 150. The input and output sections may be separate, as in the search system 100.

[0226] [Transmission lines 162 and 168] Transmission lines 162 and 168 have the function of transmitting data. Communication unit 161a, processing Data transmission and reception between the data processing unit 163a and the database 167 is via the transmission path 162. This can be done. Communication unit 161b, processing unit 163b, storage unit 165, and input / output unit 16 Data transmission and reception between points 9 can be performed via the transmission line 168.

[0227] [Processing Unit 163a and Processing Unit 163b] The processing unit 163a receives data supplied from the communication unit 161a and the database 167, etc. It has the function of performing calculations using the following. The processing unit 163b has the communication unit 161b, the storage unit 165, It also has the function of performing calculations using data supplied from the input / output unit 169, etc. Sections 163a and 163b can refer to the description of the processing unit 103. Processing unit 163a is Therefore, it is preferable that the processing capacity is higher than that of processing unit 163b.

[0228] [Storage section 165] The memory unit 165 has the function of storing the program to be executed by the processing unit 163b. The storage unit 165 stores the calculation results generated by the processing unit 163b and the data input to the communication unit 161b. It has a function to store data and other information input to the input / output unit 169.

[0229] [Database 167] Database 167 has the function of storing the search target document TD and document block TB. Furthermore, the database 167 contains the calculation results generated by the processing unit 163a, and the communication unit 161 It may have a function to store data entered into a, etc. Alternatively, server 151 In addition to the database 167, it has a separate storage unit, and this storage unit generates the data generated by the processing unit 163a. It may also have a function to store calculation results and data input to the communication unit 161a. stomach.

[0230] [Communication section 161a and communication section 161b] Using communication units 161a and 161b, data is transmitted between server 151 and terminal 152. It can send and receive data. Communication units 161a and 161b include a hub and a log A router, modem, etc. can be used. Data can be transmitted and received using either a wired connection or wireless (e.g.) For example, radio waves, infrared rays, etc. may be used.

[0231] This embodiment can be combined with other embodiments as appropriate. [Explanation of Symbols]

[0232] S1: Text, S2: Text, S3: Text, S26: Text, STB: Search text block, STD: Search Search documents, STS1: document, STS2: document, STSp: document, TB: document block, TB1: Text block, TB2: Text block, TB3: Text block, TB4: Text block, TB6: Text block, TB7: Text block, TB9: Text block, TB62: Text Block, TD: Searchable document, TD1: Searchable document, TD2: Searchable document, TDn : Document to be searched, 100: Document search system, 101: Input unit, 102: Transmission line, 103 : Processing unit, 105: Storage unit, 107: Database, 109: Output unit, 110: First pair Elephant, 110(i): First object, 120: Second object, 120(i): Second object, 15 0: Document search system, 151: Server, 152: Terminal, 161a: Communications department, 161b: Communication unit, 162: transmission line, 163a: processing unit, 163b: processing unit, 165: memory unit, 16 7: Database, 168: Transmission line, 169: Input / Output section

Claims

[Claim 1] A document search system that searches for a specific block of text similar to the search document from multiple target documents, The first step is to divide the aforementioned search document to create w (where w is a natural number of 2 or more) search document blocks, A second step is to determine the i-th search text block (where i is a natural number less than or equal to w) from the aforementioned w search text blocks, A third step involves performing a full-text search using the i-th searchable text block as a search criterion, with the multiple text blocks that are part of the multiple searchable documents being treated as a first target, thereby calculating the degree of relevance of each of the multiple text blocks included in the first target to the i-th searchable text block. A fourth step is to determine a second target from the first target that includes multiple sentence blocks with a high degree of relevance, A fifth step involves calculating the sum of the similarity scores of the sentences with the highest similarity to each sentence in the i searchable sentence block for each of the multiple sentence blocks included in the second target, and obtaining a normalized similarity score by dividing this sum by the number of sentences included in the i searchable sentence block. A sixth step in which the second to fifth steps described above are repeated for each of the w searchable text blocks, A document search system having the function of performing a seventh step after the sixth step, which is to select the document block with the highest standardization similarity as the specified document block.