Question-and-answer text retrieval apparatus and question-and-answer text retrieval program
The Q&A sentence retrieval device addresses the challenge of incomplete input by using modified cosine similarity to enhance the retrieval of relevant answers, ensuring semantic accuracy and stability.
Patent Information
- Application Number
- JP2024002871
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-07-24
- Estimated Expiration
- 2044-01-11
AI Technical Summary
Existing Q&A systems struggle to retrieve relevant answers when the input is a partial or incomplete sentence, as they rely solely on cosine similarity which cannot accurately determine the semantic meaning of incomplete character strings.
A Q&A sentence retrieval device that converts input character strings and stored question/answer text into vectors, using modified cosine similarity to prioritize matches based on word matching ratios, ensuring semantic similarity and stability in retrieval.
Stably retrieves suitable Q&A sentences even with incomplete input, improving accuracy and relevance by considering both semantic similarity and word consistency.
Smart Images

Figure 2025109137000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a Q&A text retrieval device and a Q&A text retrieval program for retrieving a Q&A text corresponding to a character string input as representing a question.
Background Art
[0002] In a field such as an operator's work at a call center where an answer is returned to an inquiry, a system that retrieves and outputs (for example, displays on a monitor) an answer text corresponding to the inquiry content is used, and a specific example thereof is described in Patent Document 1. The system described in Patent Document 1 includes a database that stores FAQs consisting of combinations of questions assumed to be inquiries and answers to those questions.
[0003] The system generates a search query by extracting keywords etc. from the conversation between the inquirer and the responder who received the question from the inquirer, and based on the search query, searches for an answer text (FAQ) stored in the database corresponding to the question from the inquirer and displays it on the monitor. By using this system, the responder can determine the answer to be returned to the inquirer by referring to the answer text displayed on the monitor. Patent Document 1 discloses that an optimal search for an answer text based on a search query is performed by keyword matching or concept search between the search query and the answer text.
[0004] Here, since both the search query and the answer text are character strings, as a method for comparing the semantic similarity of different character strings, for example, the adoption of the cosine similarity described in Patent Document 2 can be considered. The cosine similarity represents a character string as a vector in a feature space and quantitatively represents the semantic similarity of different character strings numerically from the angle formed by the vectors. Considering that the vector represents the meaning of the character string, it is important that the meaning of the character string can be uniquely defined.
Prior Art Documents
Patent Documents
[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2022-66489 [Patent Document 2] Japanese Patent Application Laid-Open No. 2023-084644 [Summary of the Invention] [Problems to be Solved by the Invention]
[0006] In this regard, for example, in the operator's work at a call center, when an operator inputs keywords for searching for an answer through a conversation with a questioner and searches for an appropriate Q&A sentence from the Q&A sentences stored in a database, that is, when searching for a Q&A sentence based on a character string that is not yet a complete sentence and whose meaning cannot be uniquely determined (for example, a character string in which three nouns are input with spaces in between), there arises a problem that a Q&A sentence corresponding to the question cannot be searched even by simply using the cosine similarity.
[0007] The present invention has been made in view of such circumstances, and an object thereof is to provide a Q&A sentence search device and a Q&A sentence search program capable of stably searching for a Q&A sentence suitable for a question even if the character string input based on the question is not a complete sentence. [Means for Solving the Problems]
[0008] The Q&A sentence retrieval device according to the first invention that meets the above object includes a text storage unit that stores a plurality of Q&A sentences each consisting of a corresponding question text sentence and answer text sentence, and a search unit that searches for the Q&A sentence corresponding to the input character string based on one or more words given as a character string representing a question. The Q&A sentence retrieval device has a vector storage unit that stores a plurality of question vectors obtained by converting a plurality of the question text sentences into vectors in a feature space, and a vector conversion unit that converts the input character string into an input vector represented by a vector in the feature space. The search unit adopts a modified cosine similarity obtained by correcting the cosine similarity as the similarity between the input vector and the question vector, and preferentially selects, as the Q&A sentence corresponding to the input character string, the Q&A sentence corresponding to the question vector with a larger modified cosine similarity. The larger the matching ratio between the words included in the input character string that is the basis of the input vector and the words included in the question text sentence that is the basis of the question vector, the larger the modified cosine similarity becomes.
[0009] The Q&A sentence retrieval device according to the second invention that meets the above object includes a text storage unit that stores a plurality of Q&A sentences each consisting of a corresponding question text sentence and answer text sentence, and a search unit that searches for the Q&A sentence corresponding to the input character string based on one or more words given as a character string representing a question. The Q&A sentence retrieval device has a vector storage unit that stores a plurality of answer vectors obtained by converting a plurality of the answer text sentences into vectors in a feature space, and a vector conversion unit that converts the input character string into an input vector represented by a vector in the feature space. The search unit adopts a modified cosine similarity obtained by correcting the cosine similarity as the similarity between the input vector and the answer vector, and preferentially selects, as the Q&A sentence corresponding to the input character string, the Q&A sentence corresponding to the answer vector with a larger modified cosine similarity. The larger the matching ratio between the words included in the input character string that is the basis of the input vector and the words included in the answer text sentence that is the basis of the answer vector, the larger the modified cosine similarity becomes.
[0010] The Q&A sentence retrieval device according to the third invention that meets the above object is a Q&A sentence retrieval device having a text storage unit that stores a plurality of Q&A sentences each consisting of a paired question text sentence and answer text sentence, and a retrieval unit that retrieves the Q&A sentence corresponding to the input character string based on an input character string consisting of one or more words given as a character string representing a question. The device includes a vector storage unit that stores a plurality of question vectors obtained by respectively converting a plurality of the question text sentences into vectors in a feature space and a plurality of answer vectors obtained by respectively converting a plurality of the answer text sentences into vectors in the feature space, and a vector conversion unit that converts the input character string into an input vector represented by a vector in the feature space. The retrieval unit respectively adopts first and second corrected cosine similarities obtained by correcting the cosine similarity as the similarity between the input vector and the question vector and the similarity between the input vector and the answer vector, and preferentially selects as the Q&A sentence corresponding to the input character string the Q&A sentence corresponding to the question vector with a larger first corrected cosine similarity and the Q&A sentence corresponding to the answer vector with a larger second corrected cosine similarity. The first corrected cosine similarity increases as the matching ratio between the words included in the input character string that is the basis of the input vector and the words included in the question text sentence that is the basis of the question vector increases. The second corrected cosine similarity increases as the matching ratio between the words included in the input character string that is the basis of the input vector and the words included in the answer text sentence that is the basis of the answer vector increases.
[0011] The Q&A text search program according to the fourth invention that meets the above object is a Q&A text search program that causes a computer to perform a process of searching for the Q&A text corresponding to the input character string from a plurality of Q&A texts each consisting of a corresponding question text sentence and answer text sentence stored in a text storage unit based on an input character string composed of one or more phrases given as a character string representing a question. The program includes a process of converting the input character string into an input vector represented as a vector in a feature space, and adopting a modified cosine similarity obtained by correcting the cosine similarity as each similarity of a plurality of question vectors obtained by converting a plurality of the question text sentences into vectors in the feature space with respect to the input vector. The program causes the computer to perform a process of preferentially selecting, as the Q&A text corresponding to the input character string, the Q&A text corresponding to the question vector with a larger modified cosine similarity. The modified cosine similarity increases as the matching ratio of the phrases included in the input character string that is the basis of the input vector and the phrases included in the question text sentence that is the basis of the question vector increases.
[0012] The Q&A text search program according to the fifth invention that meets the above object is a Q&A text search program that causes a computer to perform a process of searching for the Q&A text corresponding to the input character string from a plurality of Q&A texts each consisting of a corresponding question text sentence and answer text sentence stored in a text storage unit based on an input character string composed of one or more phrases given as a character string representing a question. The program includes a process of converting the input character string into an input vector represented as a vector in a feature space, and adopting a modified cosine similarity obtained by correcting the cosine similarity as each similarity of a plurality of answer vectors obtained by converting a plurality of the answer text sentences into vectors in the feature space with respect to the input vector. The program causes the computer to perform a process of preferentially selecting, as the Q&A text corresponding to the input character string, the Q&A text corresponding to the answer vector with a larger modified cosine similarity. The modified cosine similarity increases as the matching ratio of the phrases included in the input character string that is the basis of the input vector and the phrases included in the answer text sentence that is the basis of the answer vector increases.
[0013] The Q&A text search program according to the sixth invention that meets the above object causes a computer to perform a process of searching for the Q&A text corresponding to the input character string from a plurality of Q&A texts each consisting of a corresponding question text sentence and an answer text sentence stored in a text storage unit based on an input character string composed of one or more words given as a character string representing a question. The program includes a process of converting the input character string into an input vector expressed as a vector in a feature space, adopting, as each similarity of a plurality of question vectors obtained by converting a plurality of the question text sentences into vectors in the feature space with respect to the input vector, a first corrected cosine similarity obtained by correcting a cosine similarity, and preferentially selecting, as the Q&A text corresponding to the input character string, the Q&A text corresponding to the question vector having a larger first corrected cosine similarity; and a process of adopting, as each similarity of a plurality of answer vectors obtained by converting a plurality of the answer text sentences into vectors in the feature space with respect to the input vector, a second corrected cosine similarity obtained by correcting a cosine similarity, and preferentially selecting, as the Q&A text corresponding to the input character string, the Q&A text corresponding to the answer vector having a larger second corrected cosine similarity. The computer is made to perform these processes. The first corrected cosine similarity increases as the matching ratio of the words included in the input character string that is the basis of the input vector and the words included in the question text sentence that is the basis of the question vector becomes higher. The second corrected cosine similarity increases as the matching ratio of the words included in the input character string that is the basis of the input vector and the words included in the answer text sentence that is the basis of the answer vector becomes higher.
Effect of the Invention
[0014] The Q&A text retrieval device according to the first invention includes a vector storage unit that stores a plurality of question vectors obtained by respectively converting a plurality of question text sentences into vectors in a feature space, and a vector conversion unit that converts an input character string into an input vector expressed as a vector in the feature space. The retrieval unit adopts a modified cosine similarity obtained by correcting the cosine similarity as the similarity between the input vector and the question vector. The higher the modified cosine similarity, the more preferentially the Q&A text corresponding to the question vector is selected as the Q&A text corresponding to the input character string. Since the modified cosine similarity increases as the matching ratio between the words included in the input character string that forms the basis of the input vector and the words included in the question text sentence that forms the basis of the question vector increases, it is possible to retrieve Q&A texts that take into account the semantic similarity and word consistency as sentences. Even if the character string input based on the question is not a complete sentence, it is possible to stably retrieve a Q&A text suitable for the question.
[0015] Also, since the Q&A text retrieval device according to the second invention and the Q&A text retrieval device according to the third invention have the same configuration as the Q&A text retrieval device according to the first invention, even if the character string input based on the question is not a complete sentence, it is possible to stably retrieve a Q&A text suitable for the question.
[0016] Furthermore, since the Q&A text retrieval program according to the fourth invention, the Q&A text retrieval program according to the fifth invention, and the Q&A text retrieval program according to the sixth invention respectively correspond to the Q&A text retrieval device according to the first invention, the Q&A text retrieval device according to the second invention, and the Q&A text retrieval device according to the third invention, even if the character string input based on the question is not a complete sentence, it is possible to stably retrieve a Q&A text suitable for the question.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Embodiments for Carrying Out the Invention
[0018] Subsequently, with reference to the attached drawings, embodiments embodying the present invention will be described to facilitate understanding of the present invention. As shown in FIGS. 1 and 2, a Q&A sentence search device 10 according to a first embodiment of the present invention includes a text storage unit 14 that stores a plurality of Q&A sentences 13 each consisting of a pair of a question text sentence 11 and an answer text sentence 12, and a search unit 16 that searches for a Q&A sentence 13 corresponding to an input character string 15 based on the input character string 15 consisting of one or more words given as a character string representing a question. Details will be described below.
[0019] The Q&A sentence search device 10 is composed of software such as an operating system and a software program for Q&A sentence search (hereinafter also referred to as "Q&A sentence search program") and hardware such as a computer or a storage medium on which the software is installed. Therefore, the text storage unit 14, the search unit 16, and other means and units included in the Q&A sentence search device 10 are each composed of either one or both of software and hardware.
[0020] The computer mentioned here means, for example, a tablet-type, notebook-type, and desktop-type personal computer, a mobile terminal, and a server, and there is no limit to the number of computers the Q&A sentence search device 10 has. Therefore, the Q&A sentence search device 10 may have one computer or may have a plurality of computers. In addition, in the present embodiment, the Q&A text retrieval device 10 is used in a call center that returns an answer to a question received from a user of a product or service, but is not limited thereto. For example, the Q&A text retrieval device 10 may be used as a system in which a user of a product or service inputs a matter (question) that he or she wants to inquire about and obtains an answer thereto.
[0021] As shown in FIGS. 1 and 2, the Q&A text retrieval device 10 includes an input unit 17 to which an input character string 15 input by a user of the Q&A text retrieval device 10 (hereinafter also simply referred to as "user") using an input device (not shown) is given, and a vector conversion unit 19 capable of converting a text sentence into a vector in a feature space (high-dimensional space, vector space). Note that a known learned model is adopted as a method for encoding a text sentence into a vector.
[0022] The input device is a keyboard, a microphone, or the like. When the input device is a keyboard, the user can input the input character string 15 by operating the keyboard. When the input device is a microphone, the user can input the input character string 15 by speaking into the microphone. One input character string 15 is a text sentence composed of one or more phrases. A phrase means a word (for example, "update", "certificate") or a clause composed of a plurality of words (for example, "exceptional treatment", "certificate not submitted"). In the present embodiment, when one input character string 15 is composed of a plurality of phrases, a space for one character is inserted between the phrases in the input character string 15, and the input character string 15 is given to the input unit 17.
[0023] The vector conversion unit 19 connected to the input unit 17 can acquire the input character string 15 from the input unit 17, and based on the acquired input character string 15, generates an input vector 20 representing the input character string 15 as a vector in the feature space. That is, the vector conversion unit 19 converts the input character string 15 into the input vector 20. By vectorizing the input character string 15, for example, the semantic similarity of each of a plurality of input character strings 15 as sentences can be quantitatively handled.
[0024] The Q&A text 13 stored in the text memory unit 14 is a question text sentence 11 and an answer text sentence 12 in which a question and an answer to the question are each converted into character data. The Q&A text 13 is either prepared in advance as an assumed Q&A text before using the Q&A text retrieval device 10 or obtained during the performance of operations using the Q&A text retrieval device 10. The text memory unit 14 stores each Q&A text 13 by associating an identification number with the question text sentence 11 and the answer text sentence 12.
[0025] The retrieval means 16 is connected to the input unit 17, the vector conversion means 19, and the text memory unit 14. The retrieval means 16 can acquire the input character string 15 from the input unit 17, the input vector 20 from the vector conversion means 19, and the question text sentence 11 and the answer text sentence 12 (Q&A text 13) from the text memory unit 14, respectively. In addition, a vector memory unit 22 storing a plurality of question vectors 21 respectively corresponding to all the question text sentences 11 stored in the text memory unit 14 is connected to the retrieval means 16.
[0026] The question vector 21 is a vector obtained by converting the question text sentence 11 into a vector in the feature space by a vector conversion means (not shown), and a plurality of question vectors 21 are stored in the vector memory unit 22 in advance. The vector memory unit 22 stores each question vector 21 by associating it with an identification number common to the identification number stored by the text memory unit 14 in association with each question text sentence 11. When a new question text sentence 11 is added to the text memory unit 14 together with an identification number, the question vector 21 corresponding to the question text sentence 11 is added to the vector memory unit 22 together with the same identification number.
[0027] In the present embodiment, the vector storage unit 22 groups and stores each question vector 21 such that question vectors 21 that are similar to each other are included in the same group. As a result, the question vectors 21 stored in the vector storage unit 22 are targets for approximate nearest neighbor search, which will be described later. Assuming that there are hundreds of thousands of question vectors 21 stored in the vector storage unit 22, for example, the number of question vectors 21 included in each group is around one thousand. Also, the dimensionality of the feature space of the question vector 21 is the same as the dimensionality of the feature space of the input vector 20.
[0028] The search means 16 includes a nearest neighbor search unit 23 that acquires the input vector 20 from the vector conversion means 19, and a similarity calculation unit 24 that calculates the similarity between the input vector 20 and the question vector 21. As shown in FIG. 2, when a new input character string 15 is given from the input unit 17 to the vector conversion means 19, the vector conversion means 19 converts the input character string 15 into an input vector 20 and gives it to the nearest neighbor search unit 23 (search means 16).
[0029] Each time the nearest neighbor search unit 23 acquires one input vector 20 from the vector conversion means 19, by approximate nearest neighbor search, from all the question vectors 21 grouped and stored in the vector storage unit 22, it selects one group that contains many question vectors 21 with a high cosine similarity (not the modified cosine similarity) to the acquired input vector 20, and acquires all the question vectors 21 included in that group.
[0030] In the present embodiment, the plurality of query vectors 21 acquired by the nearest neighbor search unit 23 from the vector storage unit 22 (that is, each group includes two or more query vectors 21). Also, as approximate nearest neighbor search, in the present embodiment, a random projection tree is adopted, but it is not limited thereto, and for example, LSH (Locality Sensitive Hashing) or a KD tree may be adopted. The vector storage unit 22 stores the query vectors 21 in a format suitable for the adopted approximate nearest neighbor search algorithm.
[0031] As shown in FIG. 1, the similarity calculation unit 24 is connected to the input unit 17, the vector conversion means 19, the nearest neighbor search unit 23, and the text storage unit 14. Each time the input character string 15 is given to the input unit 17, the similarity calculation unit 24 acquires the input character string 15 from the input unit 17 directly or via a memory (storage medium). Further, each time the vector conversion means 19 converts the input character string 15 into the input vector 20, the similarity calculation unit 24 acquires the input vector 20 from the vector conversion means 19 directly or via a memory.
[0032] Also, each time the nearest neighbor search unit 23 acquires a plurality (here, T) of query vectors 21 from the vector storage unit 22, the similarity calculation unit 24 acquires the plurality of query vectors 21 from the nearest neighbor search unit 23. Further, based on the identification numbers associated with the acquired plurality of query vectors 21, the query text sentences 11 corresponding to the acquired query vectors 21 (that is, the query text sentences 11 that are the basis of the acquired query vectors 21) and the answer text sentences 12 (that is, the Q&A sentences 13 corresponding to the acquired query vectors 21) are acquired from the text storage unit 14.
[0033] The similarity calculation unit 24 calculates the similarity of the acquired input vector 20 with respect to the plurality of query vectors 21 acquired from the nearest neighbor search unit 23. Therefore, when the number of query vectors 21 stored in the vector storage unit 22 is N, the search means 16 selects, by approximate nearest neighbor search, a number of query vectors 21 less than N as the objects for calculating the modified cosine similarity.
[0034] The similarity calculation unit 24 adopts a modified cosine similarity obtained by correcting the cosine similarity as the similarity between the input vector 20 and the question vector 21. Among the question texts 13 corresponding to the question vectors 21 with a larger modified cosine similarity, the question text 13 (the question text 13 having the corresponding question vector 21) is preferentially selected as the question text 13 corresponding to the input character string 15 (treated as the question text 13 highly likely to correspond to the input character string 15). An output unit 25 capable of causing a monitor (an example of an output device), not shown in the figure, to display (output) the question text 13 is connected to the similarity calculation unit 24.
[0035] In the present embodiment, the similarity calculation unit 24 acquires T question texts 13 from the text storage unit 14, and selects only P question texts 13 corresponding to the top P question vectors 21 (T > P) with a larger modified cosine similarity among the acquired T question texts 13 as the objects to be displayed on the monitor. The similarity calculation unit 24 further sorts the selected P question texts 13 so that the question text 13 with a larger modified cosine similarity corresponding to the question text 13 is displayed (for example, displayed on the upper side of the list of question texts 13) as the question text 13 highly likely to correspond to the input character string 15 on the monitor, and gives the P question texts 13 to the output unit 25.
[0036] Here, taking θ as the angle formed by the input vector 20 and the question vector 21 in the feature space, the cosine similarity is cosθ and is only affected by the angle formed by the input vector 20 and the question vector 21. In contrast, the modified cosine similarity is affected not only by the angle formed by the input vector 20 and the question vector 21 but also by the matching ratio between the words included in the input character string 15 acquired from the input unit 17 (the input character string 15 that is the basis of the input vector 20 acquired from the nearest neighbor search unit 23) and the words included in the question text 11 that is the basis of the question vector 21 acquired from the nearest neighbor search unit 23.
[0037] Specifically, the higher the matching ratio between the phrases included in the input string 15 and the phrases included in the question text sentence 11, the greater the modified cosine similarity, and the lower the matching ratio, the smaller the modified cosine similarity. Here, the so-called "greater" or "smaller" modified cosine similarity means relative magnitude. For example, when the ratio of the phrases in the input string 15 that match the phrases in the question text sentence 11 is 100%, the modified cosine similarity is taken as the maximum value, and when the ratio is 80%, the modified cosine similarity is taken as the value obtained by subtracting a predetermined value from the maximum value. In this case, the higher the matching ratio between the phrases included in the input string 15 and the phrases included in the question text sentence 11, the greater the modified cosine similarity, and the lower the matching ratio, the smaller the modified cosine similarity (the same applies to other embodiments). In this embodiment, S calculated by the following formula 1 is used as the modified cosine similarity.
[0038] S = cos(α×θ) Formula 1
[0039] Here, let the total number of phrases included in the input string 15 be Q, the number of phrases included in the input string 15 that are also included in the question text sentence 11 be x, and k be a coefficient where 0 < k < 1. Then, α in Formula 1 is represented by the following formula 2.
[0040] α = 1 - k×(x / Q) Formula 2
[0041] Note that, as the modified cosine similarity, something other than that expressed by Formula 1 and Formula 2 may be adopted. For example, α may be α = 1.2 - k×(x / Q), α = 1 - (1 / 2)×tanh{π(x / Q)}, α = (1 / 2)×[1 - tanh{π×((x / Q) - 1)}] or α = (3 / 4) - (1 / 4)×tanh[2π×{(x / Q) - (1 / 2)}].
[0042] By adopting the modified cosine similarity as the similarity between the input vector 20 and the question vector 21, it was experimentally verified that more useful Q&A sentences 13 for the user are displayed on the monitor compared to the case of adopting the cosine similarity. Here, the words included in both the input string 15 and the question text sentence 11 may be only the words when the string of the words included in the input string 15 and the string of the words included in the question text sentence 11 are exactly the same, or in addition to the words with exactly the same string, words with different strings may also be targeted as long as the words of both represent the same object.
[0043] Also, the reason why the search means 16 limits the target for calculating the modified cosine similarity with respect to the input vector 20 to a predetermined number of question vectors 21 instead of all the question vectors 21 stored in the vector storage unit 22 is to shorten the processing time from when the search means 16 acquires the input vector 20 from the vector conversion means 19 until it gives P Q&A sentences 13 to the output unit 25. In the present embodiment, there are many question vectors 21 stored in the vector storage unit 22 (for example, hundreds of thousands), and furthermore, since the consistency of the words of the question text sentence 11 that is the basis of the question vector 21 with respect to the words of the input string 15 that is the basis of the input vector 20 is detected when calculating the modified cosine similarity, the effect of shortening the processing time is significant.
[0044] From the above description, the Q&A sentence search program used in the Q&A sentence search device 10 is a software program that causes a computer to perform a process of searching for the Q&A sentence 13 corresponding to the input string 15 from a plurality of Q&A sentences 13 stored in the text storage unit 14 based on the input string 15 given as a string representing a question. The process includes converting the input string 15 into an input vector 20, adopting the modified cosine similarity as each similarity between the input vector 20 and a plurality of question vectors 21, and preferentially selecting as the Q&A sentence 13 corresponding to the input string 15 the Q&A sentence 13 corresponding to the question vector 21 with a larger modified cosine similarity.
[0045] The Q&A text retrieval device 10 calculates the modified cosine similarity between the input vector 20 and the question vector 21, but is not limited to this. Next, with reference to FIGS. 3 and 4, a Q&A text retrieval device 30 that calculates the modified cosine similarity of the answer vector 31 obtained by converting the answer text sentence 12 into a vector in the feature space with respect to the input vector 20 will be described. In the Q&A text retrieval device 30, the same components as those in the Q&A text retrieval device 10 are denoted by the same reference numerals, and detailed description thereof will be omitted.
[0046] As shown in FIGS. 3 and 4, the Q&A text retrieval device 30 according to the second embodiment of the present invention includes an input unit 17, a vector conversion unit 19, a text storage unit 14, and an output unit 25 that are the same as those included in the Q&A text retrieval device 10. On the other hand, the search means 32 and the vector storage unit 33 included in the Q&A text retrieval device 30 are different from the search means 16 and the vector storage unit 22 of the Q&A text retrieval device 10, respectively.
[0047] The vector storage unit 33 stores the answer vector 31 instead of the question vector 21. The search means 32 includes a nearest neighbor search unit 34 that acquires the input vector 20 from the vector conversion unit 19 and the answer vector 31 from the vector storage unit 33, respectively, and a similarity calculation unit 35 that calculates the modified cosine similarity of the answer vector 31 with respect to the input vector 20.
[0048] Each time the nearest neighbor search unit 34 acquires one input vector 20 from the vector conversion unit 19, it selects one group that contains many answer vectors 31 with high cosine similarity to the acquired input vector 20 from all the answer vectors 31 grouped and stored in the vector storage unit 33 by approximate nearest neighbor search, and acquires all the answer vectors 31 included in that group.
[0049] The similarity calculation unit 35 obtains the input string 15 directly from the input unit 17 or via the memory, and obtains the input vector 20 directly from the vector conversion means 19 or via the memory. Further, the similarity calculation unit 35 obtains the T answer vectors 31 from the nearest neighbor search unit 34 that has obtained a plurality (here, T) of answer vectors 31 from the vector storage unit 22, and obtains the question text sentence 11 and the answer text sentence 12 corresponding to each obtained answer vector 31 (that is, the Q&A sentence 13 corresponding to each obtained answer vector 31) from the text storage unit 14.
[0050] After that, the similarity calculation unit 35 calculates the modified cosine similarity with respect to the obtained input vector 20 for the T answer vectors 31 obtained from the nearest neighbor search unit 34. Therefore, when the number of answer vectors 31 stored in the vector storage unit 33 is M, the search means 32 selects, by approximate nearest neighbor search, a number of answer vectors 31 less than M as the objects for which the modified cosine similarity is calculated.
[0051] The modified cosine similarity calculated by the similarity calculation unit 35 is equivalent to the modified cosine similarity represented by the above-described formulas 1 and 2, and x in formula 2 is the number of words included in the answer text sentence 12 that is the basis of the answer vector 31 among the words included in the input string 15. That is, the higher the matching ratio between the words included in the input string 15 that is the basis of the input vector 20 and the words included in the answer text sentence 12 that is the basis of the answer vector 31, the larger the modified cosine similarity becomes.
[0052] Then, the similarity calculation unit 35 selects only the P Q&A sentences 13 corresponding to the top P (T > P) answer vectors 31 with large modified cosine similarity as the objects to be displayed on the monitor, and further sorts the P Q&A sentences 13 so that the higher the modified cosine similarity, the more likely the Q&A sentence 13 is to correspond to the input string 15 on the monitor, and gives the sorted P Q&A sentences 13 to the output unit 25.
[0053] In addition, the question-and-answer text search program used in the question-and-answer text search device 30 is a software program that causes a computer to perform a process of searching for a question-and-answer text 13 corresponding to an input character string 15 from a plurality of question-and-answer texts 13 stored in the text storage unit 14 based on the input character string 15 given as a character string representing a question. The program includes a process of converting the input character string 15 into an input vector 20, and adopts a modified cosine similarity as each similarity between the input vector 20 and a plurality of answer vectors 31. The program causes the computer to perform a process of preferentially selecting, as the question-and-answer text 13 corresponding to the input character string 15, the question-and-answer text 13 corresponding to the answer vector 31 with a larger modified cosine similarity.
[0054] Also, while the question-and-answer text search device 10 calculates only the modified cosine similarity between the input vector 20 and the question vector 21, and the question-and-answer text search device 30 calculates only the modified cosine similarity between the input vector 20 and the answer vector 31, the present invention is not limited to this. With reference to FIGS. 5 and 6, a question-and-answer text search device 50 that calculates both the modified cosine similarity between the input vector 20 and the question vector 21 and the modified cosine similarity between the input vector 20 and the answer vector 31 will be described. In the question-and-answer text search device 50, the same components as those of the question-and-answer text search device 10 or the question-and-answer text search device 30 are denoted by the same reference numerals, and detailed description thereof is omitted.
[0055] As shown in FIGS. 5 and 6, the question-and-answer text search device 50 according to the third embodiment of the present invention includes an input unit 17, a vector conversion unit 19, a text storage unit 14, and an output unit 25, which are the same as those included in the question-and-answer text search device 10 and the question-and-answer text search device 30. On the other hand, the search unit 51 and the vector storage unit 52 included in the question-and-answer text search device 50 are different from the search unit 16 and the vector storage unit 22 of the question-and-answer text search device 10, and the search unit 32 and the vector storage unit 33 of the question-and-answer text search device 30, respectively.
[0056] The vector storage unit 52 stores both the question vector 21 and the answer vector 31 by associating them with identification numbers. The search means 51 includes a nearest neighbor search unit 53 that acquires the input vector 20 from the vector conversion means 19 and acquires the question vector 21 and the answer vector 31 from the vector storage unit 52, and a similarity calculation unit 54 that calculates both the similarity between the input vector 20 and the question vector 21 (hereinafter, this similarity is referred to as "similarity P") and the similarity between the input vector 20 and the answer vector 31 (hereinafter, this similarity is referred to as "similarity Q").
[0057] In the present embodiment, the first corrected cosine similarity obtained by correcting the cosine similarity is adopted as the similarity P, and the second corrected cosine similarity obtained by correcting the cosine similarity is adopted as the similarity Q. In the present embodiment, the first corrected cosine similarity is the same as the corrected cosine similarity used by the question-and-answer text examination device 10, and the second corrected cosine similarity is the same as the corrected cosine similarity used by the question-and-answer text examination device 30.
[0058] Here, the first corrected cosine similarity may increase as the matching ratio between the words included in the input string 15 that is the basis of the input vector 20 and the words included in the question text sentence 11 that is the basis of the question vector 21 increases. The second corrected cosine similarity may increase as the matching ratio between the words included in the input string 15 and the words included in the answer text sentence 12 that is the basis of the answer vector 31 increases. In the present embodiment, both the first corrected cosine similarity and the second corrected cosine similarity are represented by the above formulas 1 and 2 (that is, they are calculated by the same mathematical formula except for the object represented by x). Note that the mathematical formula for calculating the first corrected cosine similarity and the mathematical formula for calculating the second corrected cosine similarity do not have to be the same and may be slightly different. For example, the value of the coefficient k in formula 2 may be different.
[0059] The question vectors 21 are stored in the vector storage unit 52 in a grouped manner, and the answer vectors 31 are also stored in the vector storage unit 52 in a grouped manner. Each time the nearest neighbor search unit 53 obtains one input vector 20 from the vector conversion means 19, by approximate nearest neighbor search, from all the query vectors 21 stored in the vector storage unit 52, it selects one group that contains many query vectors 21 with high similarity to the obtained input vector 20, obtains all the query vectors 21 included in that group, and by approximate nearest neighbor search, from all the answer vectors 31 stored in the vector storage unit 52, it selects one group that contains many answer vectors 31 with high similarity to the obtained input vector 20, and obtains all the answer vectors 31 included in that group.
[0060] The similarity calculation unit 54 obtains the input character string 15 directly from the input unit 17 or via the memory, and obtains the input vector 20 directly from the vector conversion means 19 or via the memory. Further, the similarity calculation unit 54 obtains the T1 query vectors 21 and the T2 answer vectors 31 from the nearest neighbor search unit 53 that has obtained a plurality (here, T1) of query vectors 21 and a plurality (here, T2) of answer vectors 31 from the vector storage unit 22, and obtains the query text sentences 11 and answer text sentences 12 corresponding to each of the obtained query vectors 21, and the query text sentences 11 and answer text sentences 12 corresponding to each of the obtained answer vectors 31 from the text storage unit 14.
[0061] After that, the similarity calculation unit 54 calculates the first modified cosine similarity with respect to the obtained input vector 20 for the T1 query vectors 21 obtained from the nearest neighbor search unit 53, and calculates the second modified cosine similarity with respect to the obtained input vector 20 for the T2 answer vectors 31 obtained from the nearest neighbor search unit 53.
[0062] Therefore, when the number of question vectors 21 and the number of answer vectors 31 stored in the vector memory unit 52 are N and M respectively, the search means 51 selects, by approximate nearest neighbor search, a number of question vectors 21 less than N as the target for calculating the first modified cosine similarity, and a number of answer vectors 31 less than M as the target for calculating the second modified cosine similarity.
[0063] Then, the similarity calculation unit 54 selects, as the target for displaying only the P question-and-answer sentences 13 corresponding to the question vectors 21 and answer vectors 31 corresponding to the top P (T1 + T2 > P) having large values from among the calculated T1 first modified cosine similarities and T2 second modified cosine similarities (T1 + T2 modified cosine similarities), and further sorts the P question-and-answer sentences 13 so that the question-and-answer sentence 13 with a higher possibility of corresponding to the input character string 15 is displayed on the monitor as the first modified cosine similarity or the second modified cosine similarity is larger, and gives the sorted P question-and-answer sentences 13 to the output unit 25.
[0064] Therefore, the similarity calculation unit 54 preferentially selects the question-and-answer sentence 13 corresponding to the question vector 21 having a large first modified cosine similarity and the question-and-answer sentence 13 corresponding to the answer vector 31 having a large second modified cosine similarity as the question-and-answer sentence 13 corresponding to the input character string 15.
[0065] The question-and-answer text search program used in the question-and-answer text search device 50 is a software program that causes a computer to perform a process of searching for a question-and-answer text 13 corresponding to an input character string 15 from a plurality of question-and-answer texts 13 stored in the text storage unit 14 based on the input character string 15 given as a character string representing a question. The program includes a process of converting the input character string 15 into an input vector 20, adopting a first modified cosine similarity as each similarity between the input vector 20 and a plurality of question vectors 21, and preferentially selecting, as the question-and-answer text 13 corresponding to the input character string 15, the question-and-answer text 13 corresponding to the question vector 21 with a larger first modified cosine similarity. The program also includes adopting a second modified cosine similarity as each similarity between the input vector 20 and a plurality of answer vectors 31, and preferentially selecting, as the question-and-answer text 13 corresponding to the input character string 15, the question-and-answer text 13 corresponding to the answer vector 31 with a larger second modified cosine similarity, and causing the computer to perform these processes.
[0066] As described above, the embodiments of the present invention have been explained. However, the present invention is not limited to the above-described forms, and all changes and the like that do not deviate from the gist are within the scope of application of the present invention. For example, in the first embodiment, the search means may use all the question vectors stored in the vector storage unit as targets for calculating the modified cosine similarity. In the second embodiment, the search means may use all the answer vectors stored in the vector storage unit as targets for calculating the modified cosine similarity.
[0067] Also, in the third embodiment, the search means can use all the question vectors stored in the vector storage unit as targets for calculating the first modified cosine similarity, and use all the answer vectors stored in the vector storage unit as targets for calculating the second modified cosine similarity.
[0068] Furthermore, when storing both the query vector and the answer vector in the vector storage unit and selecting a query-and-answer sentence that takes precedence as the query-and-answer sentence corresponding to the input character string, whether to calculate the modified cosine similarity (including the first modified cosine similarity and the second modified cosine similarity) for only the query vector, only the answer vector, or both the query vector and the answer vector for the input vector may be selectable by setting.
Explanation of Signs
[0069] 10: Query-and-answer sentence search device, 11: Query text sentence, 12: Answer text sentence, 13: Query-and-answer sentence, 14: Text storage unit, 15: Input character string, 16: Search means, 17: Input unit, 19: Vector conversion means, 20: Input vector, 21: Query vector, 22: Vector storage unit, 23: Nearest neighbor search unit, 24: Similarity calculation unit, 25: Output unit, 30: Query-and-answer sentence search device, 31: Answer vector, 32: Search means, 33: Vector storage unit, 34: Nearest neighbor search unit, 35: Similarity calculation unit, 50: Query-and-answer sentence search device, 51: Search means, 52: Vector storage unit, 53: Nearest neighbor search unit, 54: Similarity calculation unit
Claims
1. A Q&A sentence search device comprising: a text storage unit that stores a plurality of Q&A sentences each consisting of a paired question text sentence and answer text sentence; and a search unit that searches for the Q&A sentence corresponding to an input character string based on one or more words given as a character string representing a question, wherein: a vector storage unit that stores a plurality of question vectors obtained by respectively converting the plurality of question text sentences into vectors in a feature space; vector conversion means for converting the input character string into an input vector represented by a vector in the feature space; the search unit employs a modified cosine similarity obtained by correcting the cosine similarity as the similarity between the input vector and the question vector, and preferentially selects, as the Q&A sentence corresponding to the input character string, the Q&A sentence corresponding to the question vector having a larger modified cosine similarity; the modified cosine similarity is characterized in that it increases as the matching ratio between the words included in the input character string that is the basis of the input vector and the words included in the question text sentence that is the basis of the question vector is higher.
2. In the Q&A sentence search device according to Claim 1, when the number of the question vectors stored in the vector storage unit is N, the search unit selects, by approximate nearest neighbor search, a number of the question vectors less than N as the objects for calculating the modified cosine similarity.
3. A Q&A sentence search device comprising: a text storage unit that stores a plurality of Q&A sentences each consisting of a paired question text sentence and answer text sentence; and a search unit that searches for the Q&A sentence corresponding to an input character string based on one or more words given as a character string representing a question, wherein: a vector storage unit that stores a plurality of answer vectors obtained by respectively converting the plurality of answer text sentences into vectors in a feature space; vector conversion means for converting the input character string into an input vector represented by a vector in the feature space; the search unit employs a modified cosine similarity obtained by correcting the cosine similarity as the similarity between the input vector and the answer vector, and preferentially selects, as the Q&A sentence corresponding to the input character string, the Q&A sentence corresponding to the answer vector having a larger modified cosine similarity; A question-and-answer text search device characterized in that the corrected cosine similarity increases as the matching ratio between the words contained in the input string that is the basis of the input vector and the words contained in the answer text sentence that is the basis of the answer vector becomes higher.
4. In the question-and-answer text search device according to claim 3, when the number of the answer vectors stored in the vector storage unit is M, the search means selects, by approximate nearest neighbor search, a number of the answer vectors less than M as objects for calculating the corrected cosine similarity. A question-and-answer text search device characterized by this.
5. A question-and-answer text search device having a text storage unit that stores a plurality of question-and-answer texts each consisting of a pair of a question text sentence and an answer text sentence, and a search means for searching for the question-and-answer text corresponding to the input string based on one or more words given as a string representing the question, A vector storage unit that stores a plurality of question vectors obtained by respectively converting a plurality of the question text sentences into vectors in a feature space and a plurality of answer vectors obtained by respectively converting a plurality of the answer text sentences into vectors in the feature space, Vector conversion means for converting the input string into an input vector represented by a vector in the feature space, The search means respectively adopts first and second corrected cosine similarities obtained by correcting the cosine similarity as the similarity between the input vector and the question vector and the similarity between the input vector and the answer vector, and the question-and-answer text corresponding to the question vector with the larger first corrected cosine similarity and the question-and-answer text corresponding to the answer vector with the larger second corrected cosine similarity are preferentially selected as the question-and-answer text corresponding to the input string. The first corrected cosine similarity increases as the matching ratio between the words contained in the input string that is the basis of the input vector and the words contained in the question text sentence that is the basis of the question vector becomes higher, and the second corrected cosine similarity increases as the matching ratio between the words contained in the input string that is the basis of the input vector and the words contained in the answer text sentence that is the basis of the answer vector becomes higher. A question-and-answer text search device characterized by this.
6. In the question-and-answer text retrieval device according to claim 5, when the number of the question vectors and the number of the answer vectors stored in the vector storage unit are N and M, respectively, the retrieval means selects, by approximate nearest neighbor search, a number of the question vectors less than N as calculation targets of the first corrected cosine similarity, and selects a number of the answer vectors less than M as calculation targets of the second corrected cosine similarity. A question-and-answer text retrieval device characterized by this.
7. A question-and-answer text retrieval program that causes a computer to perform a process of retrieving, from a plurality of question-and-answer texts each consisting of a corresponding question text sentence and an answer text sentence stored in a text storage unit, a question-and-answer text corresponding to the input character string based on an input character string consisting of one or more phrases given as a character string representing a question, a process of converting the input character string into an input vector represented by a vector in a feature space; a process of causing the computer to perform a process of adopting, as each similarity of the plurality of question vectors obtained by converting a plurality of the question text sentences into vectors in the feature space with respect to the input vector, a corrected cosine similarity obtained by correcting the cosine similarity, and preferentially selecting, as the question-and-answer text corresponding to the input character string, the question-and-answer text corresponding to the question vector having a larger corrected cosine similarity; A question-and-answer text retrieval program characterized in that the corrected cosine similarity increases as the matching ratio between the phrases included in the input character string that is the basis of the input vector and the phrases included in the question text sentence that is the basis of the question vector becomes higher.
8. A question-and-answer text retrieval program that causes a computer to perform a process of retrieving, from a plurality of question-and-answer texts each consisting of a corresponding question text sentence and an answer text sentence stored in a text storage unit, a question-and-answer text corresponding to the input character string based on an input character string consisting of one or more phrases given as a character string representing a question, a process of converting the input character string into an input vector represented by a vector in a feature space; As the respective similarities of a plurality of answer vectors obtained by converting the plurality of the answer text sentences into vectors in the feature space with respect to the input vector, a corrected cosine similarity obtained by correcting the cosine similarity is adopted, and the computer is caused to perform a process of preferentially selecting, as the answer sentence corresponding to the input character string, the answer sentence corresponding to the answer vector having a larger corrected cosine similarity. The corrected cosine similarity is characterized in that it increases as the matching ratio of the words included in the input character string that is the basis of the input vector and the words included in the answer text sentence that is the basis of the answer vector increases. An answer sentence search program.
9. An answer sentence search program for causing a computer to perform a process of searching for an answer sentence corresponding to an input character string from a plurality of answer sentences each consisting of a corresponding question text sentence and an answer text sentence stored in a text storage unit based on an input character string consisting of one or more words given as a character string representing a question. A process of converting the input character string into an input vector represented by a vector in a feature space. As the respective similarities of a plurality of question vectors obtained by converting the plurality of the question text sentences into vectors in the feature space with respect to the input vector, a first corrected cosine similarity obtained by correcting the cosine similarity is adopted, and the computer is caused to perform a process of preferentially selecting, as the answer sentence corresponding to the input character string, the answer sentence corresponding to the question vector having a larger first corrected cosine similarity. As the respective similarities of a plurality of answer vectors obtained by converting the plurality of the answer text sentences into vectors in the feature space with respect to the input vector, a second corrected cosine similarity obtained by correcting the cosine similarity is adopted, and the computer is caused to perform a process of preferentially selecting, as the answer sentence corresponding to the input character string, the answer sentence corresponding to the answer vector having a larger second corrected cosine similarity. The first corrected cosine similarity increases as the matching ratio between the words included in the input string that is the basis of the input vector and the words included in the question text sentence that is the basis of the question vector becomes higher. The second corrected cosine similarity increases as the matching ratio between the words included in the input string that is the basis of the input vector and the words included in the answer text sentence that is the basis of the answer vector becomes higher. A question-and-answer sentence search program characterized by this.
Citation Information
Patent Citations
Answer candidate proposal system and answer candidate proposal method
JP2023091791A
Search result display device, search result display method, and program
JP2022066489A
Information processing device, information processing method and program
JP2023084644A