Document retrieval program, document retrieval apparatus, and document retrieval method

The document search program enhances accuracy by fragmenting queries and sentences, calculating similarity based on relevant fragments, and determining relevant documents, thus improving the relevance of search results.

JP2025097749APending Publication Date: 2025-07-01KK TOSHIBA

Patent Information

Application Number
JP2023214107
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing document search systems inaccurately retrieve documents due to the inclusion of extra words in input queries or search sentences, leading to irrelevant results.

Method used

A document search program that divides input queries and search sentences into text fragments, calculates similarity based on selected fragments, and determines relevant documents using these similarities.

Benefits of technology

Improves search accuracy by excluding the influence of redundant expressions, ensuring results align with user intent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025097749000001_ABST
    Figure 2025097749000001_ABST
Patent Text Reader

Abstract

To provide a document retrieval program capable of improving retrieval accuracy if an input query or retrieval sentence contains extra words in document retrieval.SOLUTION: A document retrieval program of an embodiment causes a computer to realize an input query acquisition function, a retrieval sentence acquisition function, a text fragment division function, a score calculation function, and a document determination function. The input query acquisition function acquires an input query entered by a user. The retrieval sentence acquisition function acquires a retrieval sentence from a document database. The text fragment division function divides the input query and the retrieval sentence into text fragment units consisting of clauses. The score calculation function selects a text fragment to be used in calculating the similarity between the input query and the retrieval sentence, and calculates the similarity between the retrieval sentence and the input query based on the similarity of the selected text fragment.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to a document search program, a document search apparatus, and a document search method.

Background Art

[0002] As a document search system provided on the Internet, there is a QA search system that searches for and answers QA documents related to an input query from a QA set stored on a server based on an input query such as a keyword or a sentence input by a user. The QA set is a document database in which QA documents are stored. In such a QA search system, the relevance between the words of the input query and the words in each QA document is calculated using TF-IDF or Okapi BM25, and the QA documents retrieved are ranked based on the calculation results. In addition, a method has also come to be used in which deep learning technology is utilized to vectorize the sentences included in the input query and the QA documents, and the relevance between the input query and the QA documents is calculated using the similarity between the sentence vectors.

[0003] However, in the above method, all the information in the sentences included in the input query and the QA documents is used for calculating the relevance. For this reason, when the input query or the QA sentences contain words that are not originally necessary for QA search, the extra words are considered during the calculation of the relevance, and there may be a case where a QA document with low relevance to the question sentence input as the input query is retrieved.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] The problem to be solved by the present invention is to provide a document search program, a document search device, and a document search method capable of improving the search accuracy when an input query or a search sentence contains extra words in document search using the input query and the search sentence.

Means for Solving the Problem

[0006] To solve such a problem, the document search program according to the embodiment causes a computer to realize an input query acquisition function, a search sentence acquisition function, a text fragment division function, a score calculation function, and a document determination function. The input query acquisition function acquires an input query input by a user. The search sentence acquisition function acquires the search sentence from a document database that stores search target documents and search sentences assigned to each search document. The text fragment division function generates a first text fragment obtained by dividing the input query into text fragment units composed of clauses, and a second text fragment obtained by dividing the search sentence into the text fragment units. The score calculation function calculates the similarity between the first text fragment and the second text fragment, selects a text fragment to be used for calculating the similarity between the input query and the search sentence from among the first text fragment and the second text fragment based on the similarity, and calculates the similarity between the search sentence and the input query based on the similarity of the selected first text fragment and second text fragment. The document determination function determines a document to be output from among the search documents based on the similarity between the search sentence and the input query.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Embodiments for Carrying Out the Invention

[0008] Hereinafter, embodiments of a document search program, a document search device, and a document search method will be described in detail with reference to the drawings. In the following description, components having substantially the same functions and configurations are denoted by the same reference numerals, and duplicate descriptions will be made only when necessary.

[0009] (First Embodiment) FIG. 1 is a diagram showing the configuration of a document retrieval system 1 including a document retrieval apparatus 100 according to a first embodiment. The document retrieval system 1 is a computer network system that executes an interactive document retrieval for selecting document data that matches a query from a user from among document databases. As shown in FIG. 1, the document retrieval system 1 includes a document retrieval apparatus 100, a document DB (database) 200, and a client terminal 300.

[0010] The document retrieval apparatus 100 is connected to the document DB 200 and the client terminal 300 via a network or the like. The network is, for example, a LAN (Local Area Network). Note that the connection to the network may be either a wired connection or a wireless connection. Also, the network is not limited to a LAN, and may be the Internet, a public communication line, or the like.

[0011] The document DB 200 is a computer that holds a database storing a plurality of document data to be retrieved. As the document data, for example, document data in HTML format or PDF format can be used, but any other format of data may be used. In the following description, "document data" will also be simply referred to as "document". Each document is composed of a plurality of sentences, words, symbols, and the like. The type of document to be retrieved may be a QA (Question and Answer) document, a FAQ (Frequently Asked Question) document, a report, an instruction manual, or any other type of document. A QA document is a response document in which an answer to a specific question sentence is described. The QA document may also be called. Hereinafter, in the present embodiment, a document retrieval system 1 that searches for and presents a response document corresponding to a user input from among a group of QA documents stored in the document DB 200 with a QA document as a document to be retrieved will be described as an example.

[0012] In the document DB 200, short sentences (hereinafter also referred to as search sentences) for use in retrieving QA documents are stored in association with each QA document.

[0013] The client terminal 300 is a computer used by the user of the document retrieval system 1. The client terminal 300 has, as hardware, a processor, an input device, a display device, and a communication device, and functions as a user interface of the document retrieval system 1. For example, the client terminal 300 receives an input of a query by the user via the input device. Hereinafter, the query input by the user is also referred to as an input query. The input query is text data regarding a document to be retrieved. The input query can also be referred to as a search inquiry from the user. The input query may be input as a natural sentence such as an interrogative sentence, or may be input as a word such as a search word. The word may be input as a single word, or may be input as a word string including a plurality of words. The client terminal 300 transmits the input query input by the user to the document retrieval device 100 via the network.

[0014] The document retrieval device 100 functions as a server device of the document retrieval system 1. Specifically, the document retrieval device 100 receives an input query from the client terminal 300, searches for a document related to the input query from among the documents stored in the document DB 200 based on the received input query, and transmits the search result to the client terminal 300. For example, as the search result, the name of the document that matches the input query, the document data, and the search sentence are transmitted. The client terminal 300 causes the display device to display the search result received from the client terminal 300.

[0015] FIG. 2 is a diagram showing a configuration example of the document retrieval device 100. As shown in FIG. 2, the document retrieval device 100 is a computer having a processing circuit 11, a storage device 12, an input device 13, a communication device 14, and a display device 15. Data communication between the processing circuit 11, the storage device 12, the input device 13, the communication device 14, and the display device 15 is performed via a bus. The input device 13 and the display device 15 may not be provided.

[0016] The processing circuit 11 includes a processor such as a CPU (Central Processing Unit) and a memory such as a RAM (Random Access Memory). The processing circuit 11 includes an input query acquisition unit 111, a search sentence acquisition unit 112, a text fragment division unit 113, a score calculation unit 114, and a document determination unit 115. By executing a document search program, the processing circuit 11 realizes the input query function, search sentence acquisition function, text fragment division function, score calculation function, and document determination function by the above-mentioned respective units. The document search program is stored in a non-transitory computer-readable recording medium such as the storage device 12. The document search program may be implemented as a single program that describes all the functions of the above-mentioned respective units, or may be implemented as a plurality of modules divided into several functional units. Also, the above-mentioned respective units may be implemented by an integrated circuit such as an Application Specific Integrated Circuit (ASIC). In this case, it may be implemented in a single integrated circuit or individually implemented in a plurality of integrated circuits.

[0017] The storage device 12 is composed of a ROM (Read Only Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), integrated circuit memory, etc. The storage device 12 stores a document search program and the like.

[0018] The input device 13 inputs various commands from the operator. As the input device 13, a keyboard, mouse, various switches, touch pad, touch panel display, etc. can be used. The output signal from the input device 13 is supplied to the processing circuit 11.

[0019] The communication device 14 is an interface for performing data communication between the document search device 100 and an external device connected via a network. As an example, the communication device 14 performs data communication between the document DB 200 and the client terminal 300.

[0020] The display device 15 displays various information. As the display device 15, a CRT (Cathode-Ray Tube) display, a liquid crystal display, an organic EL (Electro Luminescence) display, an LED (Light-Emitting Diode) display, a plasma display, or any other display known in the art can be appropriately used. Further, the display device 15 may be a projector.

[0021] Next, the functions executed by each part of the processing circuit 11 will be described in detail. The input query acquisition unit 111 acquires the input query input by the user. At this time, the input query acquisition unit 111 receives the input query from the client terminal 300. The input query acquisition unit 111 outputs the acquired input query to the text fragment division unit 113.

[0022] The search sentence acquisition unit 112 acquires the search sentences associated with each document stored in the document database 200. For example, the search sentence acquisition unit 112 acquires one search sentence for each document stored in the document database 200. The search sentence acquisition unit 112 outputs the acquired multiple search sentences to the text fragment division unit 113.

[0023] The text fragment division unit 113 divides each of the input query and the search sentence into a plurality of text fragments. A text fragment is a unit of an element representing the meaning of a word. For example, in the document search system 1 using Japanese, a text fragment is an element obtained by further subdividing a clause included in a sentence and is composed of one or more clauses. A text fragment may be called a word.

[0024] A clause is composed of, for example, a combination of an independent word and an adjunct that follows the independent word continuously. Note that a clause may be composed of only one independent word. An independent word is a word that has a meaning as a word alone and can form a clause alone. Independent words are, for example, nouns, verbs, and adjectives. An adjunct is a word that attaches to the back of another word to form a clause and cannot form a clause alone. Adjuncts are, for example, particles and auxiliary verbs. A phrase is a collection of clauses that includes both a subject and a predicate, and is a unit larger than a clause and smaller than a sentence. Examples of phrases include main clauses and subordinate clauses. A phrase may also be called a sentence or a phrase.

[0025] A text fragment is a unit smaller than a phrase and larger than or equal to a clause. A text fragment may be composed of, for example, one clause or a plurality of consecutive clauses. Note that even when one text fragment is composed of a plurality of clauses, the text fragment does not contain a subject and a predicate at the same time. While a clause is a unit that represents a meaning including a subject and a predicate, a text fragment is a constituent element of a clause that represents a meaning. By combining a plurality of text fragments, a clause that represents a meaning is formed.

[0026] The text fragment splitting unit 113 splits the input query acquired from the input query acquisition unit 111 into text fragments and generates a group of text fragments obtained by splitting the input query. In addition, for each search sentence acquired from the search sentence acquisition unit 112, the text fragment splitting unit 113 splits the search sentence into text fragments and generates a group of text fragments obtained by splitting the search sentence. The text fragment splitting unit 113 outputs the generated group of text fragments to the score calculation unit 114. Hereinafter, the text fragment generated by splitting the input query is also called an input fragment, and the text fragment generated by splitting the search sentence is also called a search fragment.

[0027] The score calculation unit 114 calculates a similarity score for each text fragment, and uses the similarity score to calculate a similarity score with the input query for each search sentence. The similarity score is an index indicating the similarity between a plurality of texts. As the similarity score, for example, known scores such as TF-IDF between words and Okapi BM25 can be used. Also, as the similarity score, the similarity between the feature vector of the input query and the feature vector of the search document obtained by using deep learning technology may be used. Hereinafter, the similarity score is also simply referred to as similarity.

[0028] In this embodiment, the score calculation unit 114 calculates the similarity score between the text fragments included in the input query and the text fragments included in the search sentence, aggregates the calculated similarity scores for each search, and aggregates the similarity scores aggregated for each section for each search sentence.

[0029] More specifically, the score calculation unit 114 calculates a similarity score for all combinations of combining one of the input fragments and one of the search fragments. Hereinafter, the similarity score between the search fragment and the input fragment is also referred to as fragment similarity. Next, the score calculation unit 114 uses the fragment similarity to calculate a similarity score with each input section for each search section. At this time, the score calculation unit 114 selects a search fragment to be used for calculating the document similarity based on the fragment similarity, and calculates a similarity score with the input section using the fragment similarity of the selected search fragment. For example, the score calculation unit 114 selects one search fragment with a high fragment similarity with the input fragment for each input fragment. Hereinafter, the similarity score between the search section and the input section is also referred to as section similarity. Next, the score calculation unit 114 uses the section similarity of each section to calculate a similarity score with the input fragment for each search sentence. Hereinafter, the similarity score between the search sentence and the input fragment is also referred to as document similarity.

[0030] The document decision unit 115 determines, based on the document similarity of each search sentence, a document to be output to the client terminal 300 for presenting to the user as a document related to the input query from among the documents stored in the document DB 200. Hereinafter, the document determined as the document to be presented to the user is also referred to as the presented document. For example, the document decision unit 115 extracts the search sentence with the highest document similarity to the input query among the search sentences, and determines the document associated with the extracted search sentence as the presented document. The presented document may be one or more documents, and may include a plurality of documents. The document decision unit 115 outputs information regarding the presented document to the client terminal 300. The information regarding the presented document may be, for example, the document data of the presented document obtained from the document DB 200, the name of the presented document, the search sentence associated with the presented document, or a link for accessing the presented document. Also, when a plurality of documents are determined as the presented documents in descending order of document similarity, the ranking of the document similarity may be included in the information regarding the presented document.

[0031] The client terminal 300 displays the information regarding the presented document on the display device. The user can obtain the document related to the input query by checking the displayed information. Note that the information regarding the presented document may be displayed on the display device 15 of the document search apparatus 100.

[0032] (Document Search Process) Next, the operation of the document search process executed by the document search device 100 will be described. The document search device 100 starts the document search process based on the input query entered by the user being transmitted from the client terminal 300. FIG. 3 is a flowchart showing an example of the procedure of the document search process. Also, FIG. 4 is a diagram showing an example of the data flow in the document search process. Here, as an example, the case where the text fragment is a unit composed of one clause will be described. Note that the processing procedures in each of the processes described below are merely examples, and each process can be changed as appropriate as much as possible. Also, regarding the processing procedures described below, steps can be omitted, replaced, and added as appropriate according to the embodiment.

[0033] (Step S301) In the document search process, first, the input query acquisition unit 111 acquires the input query from the client terminal 300. FIG. 5 is a diagram showing an example of the input query and the search sentence. Here, the case where the input query shown in FIG. 5 is acquired will be described as an example.

[0034] (Step S302) Next, the search sentence acquisition unit 112 acquires the search sentences associated with each QA document stored in the document DB 200 from the document DB 200. Here, the case where the search sentences shown in FIG. 5 are acquired will be described as an example.

[0035] (Step S303) Next, the text fragment division unit 113 divides the acquired input query and search sentences into a plurality of text fragments.

[0036] FIG. 6 is a diagram showing the input query shown in FIG. 5 divided into text fragment units. When dividing the input query shown in FIG. 5 into text fragment units, the text fragment division unit 113 first divides the sentence "I want to transfer to the dormitory." included in the input query into clauses. Here, since the input query is composed of one clause, the input query is divided into one clause. At this time, since it is not used for similarity calculation, the punctuation marks are deleted.

[0037] After that, the text fragment splitting unit 113 splits each of the split sections into a plurality of clauses. Here, the clause "I want to transfer to the dormitory." is split into two clauses, "to the dormitory" and "I want to transfer, but". After that, the text fragment splitting unit 113 generates text fragments including one or more clauses. Here, since text fragments composed of one clause are used, two text fragments, "to the dormitory" and "I want to transfer, but", are generated as input fragments.

[0038] Similar to the input query, the text fragment splitting unit 113 splits each search sentence shown in FIG. 5 into text fragment units. FIG. 7 is a diagram showing the search sentence of the search sentence document "Q21." shown in FIG. 5 split into text fragment units. When splitting the search sentence of the search sentence document "Q21." into text fragment units, the text fragment splitting unit 113 first splits the sentence "I want to move into the dormitory at the transfer destination, so please tell me about the procedures." included in the search sentence into sections. Here, it is split into two sections, "I want to move into the dormitory at the transfer destination, so" and "Please tell me about the procedures." At this time, since it is not used for similarity calculation, the punctuation marks are deleted.

[0039] After that, the text fragment splitting unit 113 splits each of the split sections into a plurality of clauses. Here, the clause "I want to move into the dormitory at the transfer destination, so" is split into three clauses, "at the transfer destination", "to the dormitory", and "I want to move in, so", and the clause "Please tell me about the procedures." is split into two clauses, "about the procedures" and "Please tell me." After that, the text fragment splitting unit 113 generates five text fragments, "at the transfer destination", "to the dormitory", "I want to move in, so", "about the procedures", and "Please tell me.", as search fragments.

[0040] After that, the text fragment splitter 113 outputs the two generated input fragments to the score calculator 114 as a group of clauses of the input query, and outputs the five generated search fragments to the score calculator 114 as a group of clauses of the search sentence.

[0041] (Step S304) Next, the score calculator 114 calculates the similarity score between text fragments. FIG. 8 is a diagram showing a method for calculating the similarity score between the text fragments shown in FIGS. 6 and 7. As shown in FIG. 8, the score calculator 114 calculates the similarity score for each of the five search fragments with respect to each of the two input fragments. Hereinafter, the similarity score between text fragments is also referred to as the fragment similarity.

[0042] FIG. 9 is a diagram showing the calculation results of the fragment similarity shown in FIGS. 6 and 7. In FIG. 9, the clause "I want to move into the dormitory at the transfer destination" is referred to as "Clause 1", and the clause "Please tell me about the procedures" is referred to as "Clause 2". "Similarity [to the dormitory]" in FIG. 9 indicates the similarity score of each search fragment with respect to the input fragment "to the dormitory", and "Similarity [but I want to move]" indicates the similarity score of each search fragment with respect to the input fragment "but I want to move". For example, the similarity score between the search fragment "at the transfer destination" and the input fragment "to the dormitory" is "0.82".

[0043] (Step S305) Next, the score calculation unit 114 calculates a similarity score for each section of the input query with respect to each search section using the calculation result of the fragment similarity. At this time, the score calculation unit 114 first selects, for each search section, the search fragment with the highest fragment similarity among the search fragments belonging to that section as the search fragment that matches the input fragment. At this time, for each input fragment, the search fragment with the highest fragment similarity to the input fragment is selected. Also, the score calculation unit 114 selects search fragments with high fragment similarity for each input fragment so that the same search fragment is not selected, that is, so that the selected search fragments do not overlap. Then, the fragment similarities of the search fragments selected one by one for each input fragment are aggregated, and a similarity score for each input section is calculated for each search section. Hereinafter, the similarity score for each search section with respect to the input section is also referred to as the section similarity.

[0044] For example, in the example of FIG. 9, among the search fragments belonging to section 1 of the search sentence, the search fragment "to the dormitory" is selected as the search fragment with the highest fragment similarity to the input fragment "to the dormitory", and among the search fragments belonging to section 1 of the search sentence, the search fragment "I want to move in" is selected as the search fragment with the highest fragment similarity to the input fragment "I want to transfer but".

[0045] Also, among the search fragments belonging to section 2 of the search sentence, the search fragment "Please tell me" is selected as the search fragment with the highest fragment similarity to the input fragment "to the dormitory". Among the search fragments belonging to section 2 of the search sentence, the fragment similarities of the input fragment "I want to transfer but" to "Regarding the procedure" and "Please tell me" are the same, but since the search fragment "Regarding the procedure" has already been selected for the input fragment "to the dormitory", the search fragment "Please tell me" is selected.

[0046] By the above processing, in the example of FIG. 9, as the search fragments that match the input fragment "to the dormitory", "to the dormitory" in section 1 of the search sentence and "please tell me" in section 2 are selected, and as the search fragments that match the input fragment "I want to move, but", "I want to move in" in section 1 of the search sentence and "about the procedures" in section 2 are selected. The score calculation unit 114 calculates a segment similarity indicating the similarity to the input query for each search segment. FIG. 10 is a diagram showing the segment similarity calculated using the selection result shown in FIG. 9. In the example of FIG. 10, for section 2 of the search sentence, the score calculation unit 114 adds the fragment similarity "0.82" between the input fragment "to the dormitory" and the selected search fragment "please tell me" and the fragment similarity "0.82" between the input fragment "I want to move, but" and the selected search fragment "about the procedures", and calculates the total value "1.64" as the segment similarity of section 2 to the input segment. Note that instead of the total value of each fragment similarity, various statistical values such as an average value and a median value may be used to calculate the segment similarity. Similarly, for section 1 of the search sentence, the score calculation unit 114 adds the fragment similarity "1.0" between the input fragment "to the dormitory" and the selected search fragment "to the dormitory" and the fragment similarity "0.92" between the input fragment "I want to move, but" and the selected search fragment "I want to move in", and calculates the total value "1.92" as the segment similarity of section 1 to the input segment. At this time, the search fragment "at the transfer destination" included in section 1 is not used for calculating the segment similarity. Thus, when the number of text fragments included in the search segment is larger than the number of text fragments included in the input segment, search fragments that are not used for calculating the segment similarity occur. Since such search fragments do not match any input fragment, they are redundant and extra elements, and can be regarded as containing words unnecessary for calculating the similarity. Therefore, the segment similarity of section 1 becomes the segment similarity excluding the influence of the extra search fragment "at the transfer destination".

[0047] In addition, when the selected search fragments overlap, that is, when the same search fragment is selected for multiple input fragments, the section similarity is calculated for each combination when the search fragments are selected so as not to overlap, and the search fragment of the combination with the highest section similarity is selected as the matching search fragment. Note that when the selected search fragments overlap, the same search fragment may be selected repeatedly. In this case, when aggregating the section similarity, the similarity of the repeatedly selected search fragments is used multiple times.

[0048] Also, in the examples of FIGS. 5-10, the case where there is one section included in the input query has been described. However, when there are two or more sections included in the input query, the same processing as the procedure described in FIGS. 5-10 is executed for the other sections of the input query.

[0049] (Step S306) Next, the score calculation unit 114 calculates a similarity score with the input query for each search sentence using the calculation result of the section similarity. At this time, the score calculation unit 114 first selects, for each search sentence, the section with the highest section similarity among the search sections included in the search sentence as the section that matches the input section. At this time, for each input section, the search section with the highest similarity score for that section is selected. In addition, the score calculation unit 114 selects, for each input section, a search section with a high similarity score so that the same search section is not selected, that is, so that the selected search sections do not overlap. Then, the section similarities of the search sections selected one by one for each input section are aggregated, and the similarity score for the input query is calculated for each search sentence. Hereinafter, the similarity score of the search sentence with respect to the input query is also referred to as the document similarity.

[0050] For example, in the example of FIG. 10, among the search clauses belonging to the search sentence of the search document "Q21.", clause 1 is selected as the search clause with the highest similarity score to the input clause. Here, since there is one input clause included in the input query, the clause similarity "1.92" of clause 1 is used as the document similarity of the search document "Q21." with respect to the input query. When two or more input clauses are included in the input query, for each input clause, the search clause with the highest clause similarity to the input clause is selected, and the total value obtained by adding the clause similarities of the selected search clauses is used as the document similarity of the search sentence. Instead of the total value of the clause similarities, various statistical values such as the average value and median value of the clause similarities may be used as the document similarity.

[0051] Similar to the search document "Q21.", the score calculation unit 114 calculates the document similarity between the search sentence and the input query for each of the other search documents "Q22." and "Q23.". FIG. 11 is a diagram showing the document similarities calculated for each search sentence. In FIG. 11, in addition to the calculated document similarities, the search clauses used for the calculation of the document similarities and the text fragments used for the calculation of the clause similarities of the search clauses are shown. In the example of FIG. 11, the similarity score (document similarity) of the search document "Q22." with respect to the input query is "1.85". For the calculation of the document similarity of the search document "Q22.", the clause "To transfer from the dormitory" is used. For the calculation of the clause similarity of the clause "To transfer from the dormitory", the text fragments "From the dormitory" and "To transfer" are used. Also, the similarity score (document similarity) of the search document "Q23." with respect to the input query is "1.66". For the calculation of the document similarity of the search document "Q23.", the clause "Want to know about the address change procedure" is used. For the calculation of the clause similarity of the clause "To transfer from the dormitory", the text fragments "About the address change procedure" and "Want to know" are used. The score calculation unit 114 outputs the document similarities calculated for each search sentence to the document determination unit 115 as the similarity scores of the corresponding search documents.

[0052] (Step S307) Next, the document determination unit 115 determines, based on the similarity scores of the retrieved documents, the presentation documents to be presented to the user. The presentation documents are documents that match the input query. At this time, the document determination unit 115 selects search sentences with a document similarity equal to or higher than a predetermined threshold, and determines the retrieved documents corresponding to the extracted search sentences as the presentation documents to be presented to the user. Instead of extracting search sentences with a document similarity equal to or higher than a predetermined threshold, the document determination unit 115 may select a predetermined number of search sentences in descending order of document similarity. The document determination unit 115 presents the search results to the user by transmitting information about the presentation documents as the document search results to the client terminal 300. The information about the presentation documents includes, for example, the name of each document and the similarity ranking of each document. The similarity ranking is the ranking when a plurality of documents selected as presentation documents are arranged in descending order of document similarity.

[0053] For example, in the example of FIG. 11, when the threshold is "1.80", the retrieved documents "Q21." and "Q22." with a document similarity of "1.80" or higher are determined as the presentation documents, and the retrieved document "Q23." with a document similarity smaller than "1.80" is excluded from the presentation documents. Also, it is determined that the similarity ranking of the retrieved document "Q21." is "1st" and the similarity ranking of the retrieved document "Q22." is "2nd". In FIG. 11, the retrieved document "Q21." determined to be "1st" is shown in bold.

[0054] The client terminal 300 causes a display device to display the information about the presentation documents received from the document determination unit 115 as the document search results. At this time, the retrieved document with the largest document similarity may be highlighted. The user can confirm, as the search results, the name of the document with a large similarity score for the input query and the ranking of the similarity score of that document. In addition to the name of the presentation document, the document data of the presentation document, a link for accessing the presentation document, and the search sentence of the presentation document may be displayed simultaneously.

[0055] Hereinafter, the effects of the document search device 100 that executes the document search program and the document search method according to the present embodiment will be described.

[0056] The document search device 100 according to this embodiment includes an input query acquisition unit 111, a search sentence acquisition unit 112, a text fragment division unit 113, a score calculation unit 114, and a document determination unit 115. The input query acquisition unit 111 receives the input of an input query entered by the user. The search sentence acquisition unit 112 acquires a search sentence from the document database 200. The document database 200 stores search documents to be searched and search sentences assigned to the search documents. The search document is, for example, a QA document including a question sentence and an answer sentence for the question sentence. The text fragment division unit 113 generates a first text fragment (input fragment) obtained by dividing the input query in units of text fragments and a second text fragment (search fragment) obtained by dividing the search sentence in units of text fragments. A text fragment is composed of clauses. A clause includes an independent word or an independent word and an adjunct word continuous with the independent word. For example, a text fragment is composed of one clause. The score calculation unit 114 calculates the similarity (fragment similarity) between the first text fragment (input fragment) and the second text fragment (search fragment), selects a second text fragment (search fragment) to be used for calculating the similarity (document similarity) between the input query and the search sentence from among the second text fragments (search fragments) based on the similarity (fragment similarity), and calculates the similarity (document similarity) between the search sentence and the input query based on the similarity (fragment similarity) of the selected second text fragment (search fragment). The document determination unit 115 determines a document (output document) to be output from among the search documents based on the similarity (document similarity) between the search sentence and the input query.

[0057] For example, the score calculation unit 114 calculates the similarity (section similarity) between the search section and the input section by using only the similarity (fragment similarity) of the second text fragment (search fragment) belonging to the search section to be calculated that matches the first text fragment (input fragment) belonging to the input section to be calculated among the second text fragments (search fragments) belonging to the search section to be calculated, aggregates the section similarities of the search sections belonging to the same search document, and calculates the similarity (document similarity) for each search document.

[0058] With the above configuration, according to the document search device 100 according to the present embodiment, the input query and the search sentence are divided into text fragment units, and the similarity between the input query and the search sentence can be calculated by using only similar text fragments. By searching for documents that match the input query using the similarity score (document similarity) calculated in this way, even when the search sentence contains words unnecessary for calculating the similarity, the influence on the similarity calculation of redundant expressions can be eliminated, and search results that conform to the user's question intention can be returned.

[0059] Here, as a comparative example, a case where the document similarity is calculated without using the similarity calculation in units of text fragments as in the present embodiment will be described. FIG. 12 is a diagram showing the calculation result when the document similarity is calculated using all the information in the sentence without dividing the input query and the search sentence shown in FIG. 5 into text fragment units. In FIG. 12, although an input query with the meaning of "want to move into a dormitory" is input, the similarity of the search document "Q22.", which does not match the user's intention and has the meaning of "want to move out of the dormitory", is higher than that of the search document "Q21.", which matches the user's intention. Therefore, it can be seen that an answer suitable for the user's question intention cannot be provided. The reason for this is that the search sentence of the search document "Q21." contains an extra text fragment unit element of "at the transfer destination", which reduces the document similarity between the search sentence of the search document "Q21." and the input query.

[0060] On the one hand, in this embodiment, one search fragment that matches each input fragment is selected from among a plurality of search fragments belonging to the same search clause, and the clause similarity is calculated using only the similarity score (fragment similarity) of the selected search fragment. For example, when using the input query and search sentence shown in FIG. 5, as shown in FIGS. 6-11, for calculating the similarity score (clause similarity) of clause 1, only two search fragments, i.e., the search fragment "to the dormitory" that matches the input fragment "to the dormitory" and the search fragment "want to move in" that matches the input fragment "want to transfer but", are selected, and the extra search fragment "at the transfer destination" is not selected. Therefore, the similarity of the search document "Q21." that matches the user's intention is higher than that of the search document "Q22." that does not match the user's intention, and an answer suitable for the user's question intention can be provided.

[0061] Note that a text fragment may be composed of a plurality of consecutive clauses. In this case, the text fragment splitting unit 113 splits each of the input query and the search sentence into clauses, further splits each clause into phrasal units, and then generates a first text fragment (input fragment) and a second text fragment (search fragment) composed of a plurality of consecutive clauses.

[0062] Also, in this embodiment, a document search system in which the search sentence includes an extra expression in units of text fragments is described as an example, but it can also be applied to a document search system in which the input query includes an extra expression in units of text fragments. In this case, for each search fragment, one input fragment that matches is selected, and by calculating the clause similarity and document similarity between clauses using only the selected input fragments, the influence of calculating the similarity of words unnecessary for calculating the similarity included in the input query is eliminated, and search results along with the user's question intention can be returned.

[0063] Also, the same method or different methods may be used to calculate the similarity between text fragments (fragment similarity), the similarity between sections (section similarity), and the similarity between the input query and the search document (document similarity).

[0064] (Second Embodiment) The second embodiment will be described. This embodiment is a modification of the configuration of the first embodiment as follows. Regarding the same configuration, operation, and effects as those of the first embodiment, the description will be omitted. In this embodiment, the text fragments of the search sentence used to calculate the similarity between the search sentence and the input query are concatenated to create a reconstructed sentence, and the reconstructed sentence is used as the search sentence to calculate the similarity between the search sentence and the input query.

[0065] FIG. 13 is a diagram showing a configuration example of the document search device 100 according to this embodiment. As shown in FIG. 13, the processing circuit 11 further includes a candidate selection unit 116 and a sentence reconstruction unit 117. The processing circuit 11 realizes the candidate selection function and the reconstruction function by executing a document search program.

[0066] The candidate selection unit 116 acquires, for each search sentence, the search fragment used for calculating the document similarity from the score calculation unit 114. The candidate selection unit 116 selects the acquired search fragment as a candidate for use in creating a reconstructed sentence and outputs it to the sentence reconstruction unit 117.

[0067] The sentence reconstruction unit 117 reconstructs the search sentence using the search fragment acquired from the candidate selection unit 116. At this time, the sentence reconstruction unit 117 reconstructs each search sentence by concatenating the search fragments acquired from the candidate selection unit 116. Hereinafter, the reconstructed search sentence is referred to as a reconstructed sentence. The sentence reconstruction unit 117 recalculates the document similarity with the input query using the reconstructed sentence as the search sentence. The sentence reconstruction unit 117 outputs the calculation result of the document similarity to the document determination unit 115. The document determination unit 115 determines the presentation document using the document similarity acquired from the sentence reconstruction unit 117.

[0068] (Document Search Process) Next, the operation of the document search process executed by the document search device 100 according to the present embodiment will be described. FIG. 14 is a flowchart showing an example of the procedure of the document search process. FIG. 15 is a diagram showing an example of the data flow in the document search process. Since the processes of steps S1401 to S1406 in FIG. 14 are the same as the processes of steps S301 to S306 in FIG. 3, the description thereof will be omitted.

[0069] (Step S1407) In the process of step S1406, when the document similarity of each search sentence is calculated, the candidate selection unit 116 acquires the search clause used for calculating the document similarity and the search fragment used for calculating the clause similarity of the search clause. FIG. 16 is a diagram showing the search fragments acquired in the examples of FIGS. 5 to 11. For example, in the search document "Q21.", as the clause used for calculating the document similarity, "To move out of the dormitory" is acquired, and as the text fragments used for calculating the clause similarity of this clause, "To the dormitory" and "Because I want to move in" are acquired.

[0070] (Step S1407) Next, the sentence reconstruction unit 117 generates a reconstructed sentence of the search sentence using the text fragments used for calculating the clause similarity. FIG. 17 is a diagram showing the reconstructed sentence generated using the information in FIG. 16. In the example of FIG. 17, for the search document "Q21.", a reconstructed sentence "Because I want to move into the dormitory" is generated by concatenating the acquired "To the dormitory" and "Because I want to move in". As shown in FIG. 17, for a search sentence whose document similarity is smaller than a predetermined threshold, a reconstructed sentence may not be generated.

[0071] (Step S1408) Next, the sentence reconstruction unit 117 recalculates the document similarity of each search sentence using the generated reconstructed sentence as the search sentence. FIG. 18 is a diagram showing the document similarity recalculated using the reconstructed sentence in FIG. 17.

[0072] Generally, it is known that the sentences extracted from a document contain the exact meaning in the text and express the meaning of the sentence accurately when they are longer with more characters. In this embodiment, by calculating the similarity with the input query using the reconstructed text generated only with text fragments that do not use redundant expressions, the similarity between the input query and the search sentence can be determined with high accuracy.

[0073] (Third Embodiment) The third embodiment will be described. This embodiment is a modification of the configuration of the first embodiment as follows. The description of the same configuration, operation, and effects as those of the first embodiment will be omitted. In the first and second embodiments, a method of calculating a similarity score excluding the influence of redundant expressions when the search sentence contains words unnecessary for calculating the similarity was described. In this embodiment, the case where the input query contains words unnecessary for calculating the similarity will be described.

[0074] FIG. 19 is a diagram showing a configuration example of the document search apparatus 100 according to this embodiment. As shown in FIG. 19, the processing circuit 11 further includes a text fragment selection unit 118. The processing circuit 11 realizes the text fragment selection function by executing a document search program.

[0075] The text fragment selection unit 118 acquires input fragments from the text fragment splitting unit 113, extracts redundant input fragments from the acquired input fragments, and selects the input fragments excluding the redundant input fragments as useful input fragments. For example, the text fragment selection unit 118 determines whether there is a search sentence included in other input fragments for each input fragment. Then, the text fragment selection unit 118 excludes the input fragments in which there is no search sentence existing simultaneously with other input fragments, and selects the input fragments in which there is a search sentence existing simultaneously with other input fragments as input fragments effective for similarity calculation. In other words, the text fragment selection unit 118 extracts search sentences including two or more input fragments, selects the input fragments included in any of the extracted search sentences as input fragments effective for similarity calculation, and determines the input fragments not included in any of the extracted search sentences as input fragments ineffective for similarity calculation. That is, the text fragment selection unit 118 determines that the input fragments without a search sentence included with other input fragments are redundant input fragments that are not used together with other input fragments, and does not select them as useful input fragments for similarity calculation. The text fragment selection unit 118 outputs the input fragments selected as useful input fragments for similarity calculation to the score calculation unit 114.

[0076] The score calculation unit 114 calculates the document similarity in the same manner as in the above embodiment using only the input fragments acquired from the text fragment selection unit 118. Thereby, the document similarity is calculated using only the useful input fragments excluding the redundant input fragments included in the input query.

[0077] (Document Search Processing) Next, the operation of the document search process executed by the document search device 100 according to this embodiment will be described. FIG. 20 is a flowchart showing an example of the procedure of the document search process. Also, FIG. 21 is a diagram showing an example of the data flow in the document search process. Since the processes of steps S2001 to S2003 in FIG. 20 are the same as the processes of steps S301 to S303 in FIG. 3, the description thereof will be omitted.

[0078] (Step S2004) In the process of step S2003, when the input query and each search sentence are split into text fragment units and input fragments and search fragments are generated, the text fragment selection unit 118 extracts, for each combination of two input fragments included in the same clause, the search sentences including both input fragments, and calculates the number of search sentences extracted as the co-occurrence document number.

[0079] FIG. 22 is a diagram showing an example of an input query including redundant expressions. FIG. 22 also shows the input clauses obtained by splitting the input query into clause units and the split fragments obtained by splitting the input clauses into text fragment units. Here, the input fragment "after marriage" is composed of two clauses "after marriage" and "as an opportunity", and the input fragments "to company housing" and "I want to move" are composed of one clause. Thus, the text fragments may be composed of different numbers of clauses according to the meaning and content of the characters.

[0080] FIG. 23 is an example of a search sentence obtained from the document DB 200. FIG. 24 is a diagram showing an example of the co-occurrence number calculated using the input fragments shown in FIG. 22. In FIG. 24, the co-occurrence numbers are calculated for three combinations of two out of the three input fragments "after marriage", "to company housing", and "I want to move".

[0081] Next, the text fragment selection unit 118 extracts input fragments that do not co-occur with other input fragments using the calculation result of the co-occurrence count. In an example of FIG. 24, since the co-occurrence counts of all combinations including the input fragment "after marriage" are all "0", the input fragment "after marriage" is extracted as an input fragment that does not co-occur with other input fragments. Then, the text fragment selection unit 118 determines the input fragment "after marriage" as a redundant element, and selects the input fragments "to company housing" and "I want to move, but" excluding the input fragment "after marriage" as input fragments useful for the similarity calculation.

[0082] (Step S2005 - Step S2008) Next, the score calculation unit 114 calculates the document similarity using only the input fragments selected as useful input fragments. Then, the document determination unit 115 determines the presentation document using the calculation result of the document similarity and outputs it to the client terminal 300. In the examples of FIGS. 22 to 24, since the similarity between the search sentences of the search documents "Q31." and "Q32." and the input fragments "to company housing" and "I want to move, but" is high, the document similarity of the search documents "Q31." and "Q32." becomes high and they are selected as the presentation documents. On the other hand, since the similarity between the search sentences of the search documents "Q33." and "Q34." and the input fragment "after marriage" redundantly added to the input query is high, the document similarity of the search documents "Q33." and "Q34." becomes low and they are not selected as the presentation documents.

[0083] On the other hand, when calculating the document similarity using the entire input query including the input fragment "after marriage", not only the search documents "Q31." and "Q32." with high similarity to the input fragments "to company housing" and "I want to move, but", but also the document similarity of the search documents "Q33." and "Q34." with high similarity to the input fragment "after marriage" becomes high, and search sentences that do not match the user's search intention are also extracted.

[0084] In this embodiment, by determining redundant text fragments in the input query and calculating the document similarity using only the useful input fragments obtained by excluding the redundant input fragments included in the input query, it is possible to accurately calculate the similarity of the search sentence even when there are redundant clauses in the input query. That is, even if the input query contains redundant expressions, by excluding the redundant clauses during document search, the influence on the similarity calculation can be eliminated, and search results corresponding to the user's question intention can be returned.

[0085] (The Fourth Embodiment) The fourth embodiment will be described. This embodiment is a modification of the configuration of the first embodiment as follows. Regarding the same configuration, operation, and effects as those of the first embodiment, the description will be omitted. In this embodiment, the text fragments of the search sentence that were not used for calculating the similarity with the input query are recommended as keywords when performing a search to narrow down the search results.

[0086] FIG. 25 is a diagram showing a configuration example of the document search apparatus 100 according to this embodiment. As shown in FIG. 25, the processing circuit 11 further includes a text fragment extraction unit 119 and a search word recommendation unit 120. The processing circuit 11 realizes a text fragment extraction function and a search word recommendation function by executing a document search program.

[0087] The text fragment extraction unit 119 extracts search fragments that are not similar to the text fragments of the input sentence from among the search fragments included in the search sentence of the search document determined as the presented document. Specifically, the text fragment extraction unit 119 extracts the search fragments that were not used for calculating the document similarity from among the search fragments included in the search sentence of the search document determined as the presented document, and outputs them to the search word recommendation unit 120.

[0088] The search word recommendation unit 120 outputs the search fragments obtained from the text fragment extraction unit 119 to the client terminal 300 as recommended words effective for the narrowing-down search. The narrowing-down search is performed to further narrow down a plurality of search documents presented as presentation documents that match the input query. Generally, when a plurality of search results are output, the user will perform a narrowing-down search until the desired answer is obtained. The recommended words may be called suggestion words.

[0089] In addition to the information regarding the plurality of presentation documents displayed as search results, the client terminal 300 causes the display device to display the recommended words. Further, the client terminal 300 accepts the selection of the recommended words by the user via the input device, and transmits the recommended words input by the user to the document search device 100 via the network. The document search device 100 re-searches for documents that match the input query using the recommended words received from the client terminal 300 as the input query, and transmits the search results to the client terminal 300. The client terminal 300 re-displays the search results received from the client terminal 300 on the display device.

[0090] (Document Search Process) Next, the operation of the document search process executed by the document search device 100 according to the present embodiment will be described. FIG. 26 is a flowchart showing an example of the procedure of the document search process. Further, FIG. 27 is a diagram showing an example of the data flow in the document search process. The processes of steps S2601 - S2607 in FIG. 26 are the same as the processes of steps S301 - S307 in FIG. 3, and thus the description thereof will be omitted.

[0091] (Step S2608) In the process of step S2607, when the presentation document is determined from the search documents, the text fragment extraction unit 119 extracts the search fragments that were not used for calculating the document similarity of the presentation document. FIG. 28 is an example of input fragments obtained by dividing the input query acquired from the document DB 200 into text fragment units. FIG. 29 is an example of a search sentence acquired from the document DB 200. FIG. 30 is a diagram showing the presentation document extracted from the search sentence of FIG. 29 using the input query of FIG. 28. In FIG. 30, the search documents "Q41.", "Q42.", "Q43." are extracted as the presentation documents that match the input query "I'm retiring next month." Also, in the example of FIG. 30, as the text fragments not used for the similarity calculation with the input query, the search fragment "Regarding the employment extension system" of the search document "Q41.", the search fragments "Of the defined contribution pension", "Handling" of the search document "Q42.", and the search fragments "Regarding taxes", "Procedures" of the search document "Q43." are extracted. Here, text fragments composed of general verbs or interrogative words, such as "Tell me" in the search document "Q41." and "What will happen?" in the search document "Q42.", are not suitable for the narrowing-down search words and may be excluded from the extraction target.

[0092] (Step S2609) Next, the search word recommendation unit 120 presents recommended words to the user by sending the extracted search fragments to the client terminal 300 as recommended words. The client terminal 300 causes the recommended word display device to display the recommended words received from the search word recommendation unit 120. At this time, it is advisable to remove auxiliary words and auxiliary verbs with low necessity from the word sequence of the recommended words before display. For example, in the example of FIG. 30, from the extracted recommended words "Regarding the employment extension system", "Regarding taxes", "Of the defined contribution pension", "Handling is", "Regarding procedures", the auxiliary words and auxiliary verbs are removed, and five word sequences of "Employment extension system", "Defined contribution pension", "Handling", "Taxes", and "Procedures" may be displayed as recommended words. For example, a sentence such as "There are 3 search results. You can narrow it down by 'Employment extension system', 'Defined contribution pension handling', and 'Taxes procedures.'" may be displayed on the display device.

[0093] When one of the recommended words is selected by the user, the document search device 100 performs a narrowing search by re-executing the document search process using the selected recommended word as an input query. After that, the document search device 100 presents the search results of the narrowing search to the user again. The user repeats the operation of selecting one recommended word from the plurality of presented recommended words until the target document is found.

[0094] In this embodiment, among the text fragments included in the search sentence, text fragments that were not used for calculating the similarity with the input query can be used as recommended words for the following narrowing-down search. By selecting a search word to be used in the narrowing-down search from the recommended words, the user can execute the narrowing-down search without thinking of a search word by themselves. Thereby, it is possible to assist in narrowing down the search results and improve the efficiency of narrowing down the search results. Also, when the query input by the user is short and there is insufficient information in the input query to search for appropriate documents, the part corresponding to the insufficient information in the search sentence is not used for calculating the similarity because there is no similar input fragment. Therefore, by presenting text fragments that were not used for calculating the similarity in the search sentence as candidates for the insufficient information in the input query, it is possible to prompt the input of the insufficient information.

[0095] (Fifth Embodiment) The fifth embodiment will be described. This embodiment is a modification of the configuration of the first embodiment as follows. Regarding the same configuration, operation, and effects as the first embodiment, the description will be omitted. In this embodiment, when displaying the search results, text fragments determined to be similar to the input query in the search sentence of the search document are highlighted.

[0096] FIG. 31 is a diagram showing a configuration example of the document search device 100 according to this embodiment. As shown in FIG. 31, the processing circuit 11 further includes a text fragment highlighting unit 121. The processing circuit 11 realizes the text fragment highlighting function by executing a document search program.

[0097] The text fragment highlighting unit 121 extracts a search fragment similar to the text fragment of the input query from among the search fragments included in the search query of the search document determined as the presentation document, and highlights the extracted search fragment. Specifically, the text fragment highlighting unit 121 extracts the search fragment used for calculating the document similarity, and outputs the extracted search fragment to the client terminal 300 as the text fragment to be highlighted.

[0098] When the client terminal 300 displays information regarding the presentation document on the display device, the client terminal 300 highlights and displays the search query of the presentation document with the text fragment received from the text fragment highlighting unit 121.

[0099] (Document Search Process) Next, the operation of the document search process executed by the document search apparatus 100 according to the present embodiment will be described. FIG. 32 is a flowchart showing an example of the procedure of the document search process. FIG. 33 is a diagram showing an example of the data flow in the document search process. The processes of step S3201-step S2607 in FIG. 32 are the same as the processes of step S301-step S307 in FIG. 3, and thus the description thereof will be omitted.

[0100] (Step S3208) In the process of the process of step S3207, when the presentation document is determined from the search documents, the text fragment highlighting unit 121 extracts the search fragments used for the calculation of the document similarity of the presentation document. FIG. 34 is a diagram showing the document similarity of the search documents "Q41.", "Q42.", "Q43." used in the example of FIGS. 5-11 and the text fragments used for the calculation of the document similarity. In FIG. 34, in the search documents "Q41." and "Q42." selected as the presentation documents, the section similarity of the sections used for the document similarity is shown in bold. The text fragment highlighting unit 121 extracts, as text fragments to be highlighted, the four search fragments "to the dormitory", "because I want to move in", "from the dormitory", and "to transfer" that are the text fragments used for the similarity calculation of the document similarity among the search fragments of the search documents "Q41." and "Q42." selected as the presentation documents.

[0101] (Step S3209) Next, in the same manner as in the first embodiment, the document determination unit 115 presents the search results to the user by transmitting the information regarding the presentation document to the client terminal 300. At this time, the text fragment highlighting unit 121 outputs the search fragments extracted in the process of step S3208 to the client terminal 300 as text fragments to be highlighted.

[0102] The client terminal 300 causes the display device to display the information regarding the presentation document received from the document determination unit 115. At this time, the client terminal 300 displays the search sentence of the presentation document together with the name of the presentation document. At this time, the client terminal 300 highlights the text fragments received from the text fragment highlighting unit 121 and displays the search sentence. FIG. 35 is a diagram showing an example of a display screen to be displayed on the display device using the example of FIG. 34. In the example of FIG. 35, among the search sentences of the search documents "Q41." and "Q42.", the portions corresponding to the text fragments "to the dormitory", "because I want to move in", "from the dormitory", and "to transfer" are highlighted in bold. Note that the method of highlighting is not limited to bold, and the color, size, and background of the characters may be changed, or the characters may be made to blink.

[0103] In this embodiment, among the search sentences of the documents in the search results, the search fragment used for calculating the document similarity is determined to be similar to the input query, and a search sentence highlighting the text fragment determined to be similar to the input query can be displayed. By checking the part of the search sentence determined to match the input query, the user can determine whether the search results are related to the content the user wants to know. For example, in the example of FIG. 35, the user can grasp that the content of the search document "Q41." is related to "move-in to the dormitory" and the content of the search document "Q42." is related to "move-out from the dormitory" by checking the highlighted parts in the search sentences of the search documents "Q41." and "Q42.", which makes it easier to determine which document is the one the user is looking for.

[0104] Thus, according to any of the above-described embodiments, it is possible to provide a document search apparatus, a document search method, and a document search program that can improve the search accuracy in document search using an input query and a search sentence when the input query or the search sentence contains extra words.

[0105] Note that the present invention is not limited to the above-described embodiments as they are, and at the implementation stage, the components can be modified and embodied without departing from the gist thereof. Also, various inventions can be formed by appropriately combining a plurality of components disclosed in the above-described embodiments. For example, some components may be deleted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined.

Explanation of Reference Numerals

[0106] 1…Document search system, 100…Document search device, 200…Document database, 300…Client terminal, 11…Processing circuit, 12…Storage device, 13…Input device, 14…Communication device, 15…Display device, 111…Input query acquisition unit, 112…Search sentence acquisition unit, 113…Text fragment division unit, 114…Score calculation unit, 115…Document determination unit, 116…Candidate selection unit, 117…Sentence reconstruction unit, 118…Text fragment selection unit, 119…Text fragment extraction unit, 120…Search word recommendation unit, 121…Text fragment highlighting unit.

Claims

1. A computer, an input query acquisition function for acquiring an input query input by a user, a search query acquisition function for acquiring the search query from a document database storing search target documents and search queries assigned to each search document, a text fragment splitting function for generating a first text fragment obtained by splitting the input query in units of text fragments composed of clauses, and a second text fragment obtained by splitting the search query in units of the text fragments, a score calculation function for calculating the similarity between the first text fragment and the second text fragment, selecting text fragments to be used for calculating the similarity between the input query and the search query from among the first text fragment and the second text fragment based on the similarity, and calculating the similarity between the search query and the input query based on the similarity of the selected first text fragment and second text fragment, a document determination function for determining a document to be output from among the search documents based on the similarity between the search query and the input query, A document search program for realizing the above.

2. The score calculation function calculates the similarity between the input query for each clause of the search query based on the similarity between the first text fragment and the second text fragment, and calculates the similarity between the search query and the input query based on the similarity for each clause of the search query. The document search program according to Claim 1.

3. The score calculation function selects, for each of the first text fragments, a second text fragment similar to the first text fragment based on the similarity between the first text fragment and the second text fragment, and calculates the similarity between the search query and the input query using only the selected second text fragments. The document search program according to Claim 1.

4. The search query includes words unnecessary for calculating the similarity. The document search program according to Claim 3.

5. The score calculation function selects, for each of the first text fragments, the second text fragment similar to the first text fragment so that the second text fragments do not overlap. The document search program according to Claim 3.

6. The text fragment splitting function splits each of the input query and the search sentence into clauses, and generates the first text fragment and the second text fragment using the split clauses. The document search program according to claim 1.

7. A program for further implementing a sentence reconstruction function that obtains a second text fragment similar to the first text fragment and generates a reconstructed sentence using only the obtained second text fragment. The document determination function recomputes the similarity between the search sentence and the input query using the reconstructed sentence as the search sentence. The document search program according to claim 3.

8. A program for further implementing a text fragment selection function that extracts a search sentence including a plurality of the first text fragments and selects a first text fragment included in at least one of the extracted search sentences. The score calculation function calculates a similarity score with the text fragment of the selected user question sentence and aggregates it in units of search clauses. The document search program according to claim 1.

9. The input query includes words unnecessary for similarity calculation. The document search program according to claim 8.

10. A text fragment extraction function that extracts a second text fragment that is not similar to the first text fragment from the second text fragments, and A search word recommendation function that outputs the extracted second text fragment as a recommended word for search. The document search program according to claim 3 for further implementing the functions.

11. A text fragment highlighting function that extracts a second text fragment similar to the first text fragment from the second text fragments and highlights the extracted second text fragment, for further implementing the function. The document search program according to claim 3.

12. An input query acquisition unit that acquires an input query input by a user, and A search sentence acquisition unit that acquires the search sentence from a document database that stores search target search documents and search sentences assigned to each search document. A first text fragment obtained by splitting the input query into text fragment units each composed of a clause, and a text fragment splitter that generates a second text fragment obtained by splitting the search sentence into the text fragment units, A score calculator that calculates the similarity between the first text fragment and the second text fragment, selects a text fragment to be used for calculating the similarity between the input query and the search sentence from among the first text fragment and the second text fragment based on the similarity, and calculates the similarity between the search sentence and the input query based on the similarity of the selected first text fragment and second text fragment, A document determination unit that determines a document to be output from among the search documents based on the similarity between the search sentence and the input query, A document search device comprising the above.

13. Obtaining an input query input by a user, Obtaining the search sentence from a document database that stores search documents to be searched and search sentences assigned to each search document, Generating a first text fragment obtained by splitting the input query into text fragment units each composed of a clause, and a second text fragment obtained by splitting the search sentence into the text fragment units, Calculating the similarity between the first text fragment and the second text fragment, selecting a text fragment to be used for calculating the similarity between the input query and the search sentence from among the first text fragment and the second text fragment based on the similarity, and calculating the similarity between the search sentence and the input query based on the similarity of the selected first text fragment and second text fragment, Determining a document to be output from among the search documents based on the similarity between the search sentence and the input query, A document search method comprising the above.

Citation Information

Patent Citations

  • Dialog system, dialog method, program, and storage medium

    JP2020123131A

  • Classification method, device, and program

    JP7131130B2

Cited By

  • Similar text retrieval method and system

    CN121833913A

  • Information processing system, information processing method, and information processing program

    JP7875576B1