Search formula creation apparatus and search formula creation method
The search query creation device addresses the challenge of creating effective search queries in document retrieval systems by using feedback-based refinement of search queries, resulting in improved accuracy and reduced user effort.
Patent Information
- Application Number
- JP2023191502
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-09
- Publication Date
- 2025-05-21
AI Technical Summary
Existing document retrieval systems face challenges in creating appropriate search queries, particularly for inexperienced users, as they often require trial-and-error adjustments to refine search queries and ensure relevant document sets are retrieved.
A search query creation device that assists in generating and refining search queries by creating search query correction candidates based on the evaluation of initial document sets, allowing users to provide feedback on differences in search results to identify and refine the most appropriate search queries.
This approach supports the creation of more appropriate search queries, reducing user burden and improving the accuracy of document retrieval by iteratively refining search queries based on user feedback and document set evaluations.
Smart Images

Figure 2025079075000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a technique for assisting document retrieval. [Background technology]
[0002] Currently, there are several methods for searching document databases. One is to input a keyword and search for text by keyword matching, and another is to input text, create a vector, and search documents in the document database based on their similarity to the vectorized text.
[0003] The former method is generally called keyword search or full-text search, and is widely used. In particular, in patent searches, full-text searches using keyword logical expressions (search logical expressions) are common. Search systems that employ keyword search generally do not take synonyms into account for the keywords entered by the user, so for example, even if you enter "mado" (window) as a search keyword and perform a search, documents that use the synonym "window" as the same as "mado" (window) will not be found.
[0004] The latter method is generally called vector search, and its vectorization methods include TF-IDF, BM-25, BERT, etc. In this search method, the input keyword or input text does not necessarily have to match the keywords in the search target document, but the similarity between the input text and the search target document is measured and displayed in order of similarity. For this reason, vector search tends to have higher search accuracy than keyword search.
[0005] However, in the above-mentioned vector-based search method, it is difficult to obtain evidence of why a document was searched for the input text, so there is still a high demand for full-text search using a search formula. On the other hand, when a user creates a search formula from the input text, it is easy to overlook the selection of keywords and synonyms, and a trial-and-error search is required, which is difficult. Regarding the above-mentioned search, Patent Document 1 has been proposed. Patent Document 1 discloses a technology for breaking down the input text into constituent units and creating a search formula for each constituent unit based on keywords contained in the constituent units as a search query. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] International Publication No. 2021 / 245814 Summary of the Invention [Problem to be solved by the invention]
[0007] In Patent Document 1, a user can edit a search formula generated from an input text by adding keywords or changing the search range. In this way, document searches are often performed repeatedly and by trial and error while changing the search formula. When performing searches by trial and error, the initial search formula is often rough and needs to be refined to a more appropriate search formula.
[0008] However, when a search query is changed, it is not necessarily changed so that the set of documents that the user intended is retrieved. For example, in the case of an inexperienced user, the intended document may be excluded from the set of documents that is the search result as a result of changing the search query. For this reason, when changing the search query, it is desirable to change it to a more appropriate one, that is, to search for a set of documents that is more likely to include the desired document. However, this point is not taken into consideration in Patent Document 1.
[0009] In addition, when changing the search query to a more appropriate one in this way, it can also be realized by, for example, checking how the search results (document set) have changed before and after, that is, whether the search results are approaching the intended document set. However, in Patent Document 1, the generated search query is only displayed as the input text, and it is not possible to understand what document set will be searched. As a result, the user has to check the list of retrieved documents without being able to judge the validity of the search query, and modifying the search query places a heavy burden on the user. Therefore, the present invention aims to support the creation of more appropriate search queries in document retrieval. [Means for solving the problem]
[0010] In the present invention, a search query correction candidate that corrects or changes the search query is created according to the evaluation result of the first document set with the search query created from the input text, a difference set showing the difference between the first document set and the second document set with the search query correction candidate is presented, and a corrected search query is identified from the search query correction candidate according to the confirmation result (feedback) of the presented difference set. Note that the "difference" in the present invention may be anything that indicates the difference, and may also include other expressions such as change.
[0011] More preferably, the document set for each logical sum corresponding to the keywords included in the search query or the search query modification candidates is illustrated. In addition, in the present invention, the evaluation result (scoring) can be determined based on the user's evaluation of the search query or the similarity obtained when a vector search, particularly a dense search, is performed.
[0012] More specifically, the search query creation device assists in the creation of a search query for searching documents based on input text, and includes a search query creation unit that creates a search query using keywords extracted from the input text, a search execution unit that searches for documents using the search query to create a first document set, a search query correction candidate creation unit that modifies the search query to create a search query correction candidate in accordance with an evaluation result for the first document set, and a feedback unit that receives feedback that is a confirmation result for the difference set presented to a user of the search query creation device and identifies a revised search query from the search query correction candidates in accordance with the confirmation result.
[0013] The present invention also includes a search query creation method using the search query creation device, a program for causing the search query creation device to function as a computer, and a storage medium storing the program. Furthermore, the present invention also includes a search using the search query creation device. Effect of the Invention
[0014] According to the present invention, it is possible to support the creation of a more appropriate search query for document search while reducing the burden on the user. [Brief description of the drawings]
[0015] [Figure 1] FIG. 1 is a functional block diagram of a search query creation device 100 in first to fourth embodiments. [Diagram 2] FIG. 2 is a diagram showing a search query creation screen 2000 in the first to fourth embodiments of the present invention. [Diagram 3] 1 is a flowchart showing a processing procedure in the first to fourth embodiments of the present invention. [Figure 4] 1 is a flowchart showing the procedure of a search query creation process in the first to fourth embodiments of the present invention. [Diagram 5] FIG. 1 is a diagram showing an example of a search formula DB 122 in the first to fourth embodiments of the present invention. [Figure 6]11 is a flowchart showing a processing procedure for generating search query correction candidates in the first embodiment of the present invention. [Figure 7] 1 is a flowchart showing a processing procedure of a feedback process in the first to fourth embodiments of the present invention. [Figure 8] 4 is a flowchart showing a procedure of a feedback process in the first to fourth embodiments of the present invention. [Figure 9] 4 is a flowchart showing a procedure of a feedback process in the first to fourth embodiments of the present invention. [Figure 10] 1 is a table showing documents in a difference set in Examples 1 to 4 of the present invention. [Figure 11] 1 shows an example of a feedback screen (interface) in Examples 1 to 4 of the present invention. [Figure 12] FIG. 1 is a diagram showing a schematic diagram of the expansion of a search set in Examples 1 to 4 of the present invention. [Figure 13] 13 is a visualization example in which addition or deletion of a search logical sum is presented in the process of generating a search query correction candidate in the first to fourth embodiments of the present invention. [Figure 14] 1 is a flowchart showing a procedure for generating search query correction candidates in the first to fourth embodiments of the present invention. [Figure 15] 13 is a flowchart showing a processing procedure for generating search query correction candidates in the second embodiment of the present invention. [Figure 16] 13 is a flowchart showing a processing procedure for generating search query correction candidates in the second embodiment of the present invention. [Figure 17] 13 is a flowchart showing a processing procedure for generating search query correction candidates in the third embodiment of the present invention. [Figure 18] 13 is a flowchart showing a processing procedure of a search query correction candidate generating unit in the third embodiment of the present invention. [Figure 19] 13 is a flowchart showing a processing procedure of a search query correction candidate generating unit in Example 4 of the present invention. [Figure 20] FIG. 13 is a system configuration diagram showing an example of a document search system according to a fifth embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] Before describing the contents of this embodiment, the vocabulary used in this specification will be described below. Search expression: A logical sum or logical product of one or more keywords. In this specification, logical sum is represented by "+" and logical product is represented by "x". For example, a search expression is represented as follows: "(Noise + Acoustic) x (Removal + Filter + Reduction)" Documents searched using this search formula contain "noise" or "acoustic" and also contain "removal" or "filter" or "reduction." For example, the following documents are searched: Document 1: "Provide earphones capable of eliminating noise." Document 2: "Provide flooring that can reduce noise." Nearby search: A search method in which the distance between two keywords is less than a threshold. The distance between keywords is not specified as a character unit, word unit, or phrase unit. For example, a nearby search expression is expressed in the following format: "{(noise + noise) × (remove + filter + reduce), 20n}" In the above search formula, the search criteria are that "noise" or "acoustic noise" and "remove" or "filter" or "reduce" appear within 20 characters.
[0017] The present embodiment will be described below. In this embodiment, a search query is created according to an input text input by a user, and a full-text search is performed on a target document. The present invention is not limited to full-text searches and can be applied to various searches. In this embodiment, a search query creation device is used, but the corresponding function may be provided in a search device. The search query creation support device of this embodiment has the following configuration. In a search query creation device that supports the creation of a search query for searching documents based on an input text, the search query creation device includes a search query creation unit that creates a search query using a keyword extracted from the input text, a search execution unit that searches documents using the search query to create a first document set, a search query correction candidate creation unit that corrects the search query and creates a search query correction candidate according to an evaluation result for the first document set, and a difference set that indicates the difference between the first document set and a second document set created using the search query correction candidate, and a feedback unit that accepts feedback that is a confirmation result for the difference set presented to a user of the search query creation device and identifies a corrected search query from the search query correction candidate according to the confirmation result. In this embodiment and in each of the examples described later, the difference set and the like are presented to the user, but this presentation can be omitted. In the following, each of the examples that are specific examples of this embodiment will be described. EXAMPLES
[0018] Below, Examples 1 to 5 are explained, but the creation of search query correction candidates differs between Examples 1 to 4. Also, Example 5 explains an implementation example that realizes each Example. Among these, in Example 1, a search query is created using the creation result of a search query before correction.
[0019] First, the configuration of the first embodiment will be described, which has many common parts with the second to fourth embodiments described later. Therefore, the configuration will be described collectively in the first embodiment. FIG. 1 is a functional block diagram of a search query creation device 100 in the first to fourth embodiments. In FIG. 1, the search query creation device creates a search query for an input text 1000. To this end, the search query creation device 100 has a processing unit 110 and a storage unit 120.
[0020] Processing unit 110 has a search query creation unit 111, a search execution unit 112, a search information visualization unit 113, a search query correction candidate creation unit 114, a feedback unit 115, and a document list output unit 116. Storage unit 120 has a document DB 121, a search query DB 122, a search document DB for each logical sum 123, and a correct answer document DB 124.
[0021] The specific processing performed by each unit will be described in detail using the processing flow charts in each embodiment, so the description of each component in the block diagram of Fig. 1 will be brief below. Note that the search query creation device 100 further includes or is connected to an input unit and an output unit (not shown). These are used to accept user operations and present various information.
[0022] First, the search formula creation unit 111 creates a search formula using the input text 1000 input by the user via an input unit (not shown). This creation method includes various methods. For example, important words are extracted from the input text 1000, and a search formula is created using the extracted important words as keywords. The input text may be treated as a word unit and used as a keyword to create a search formula. In these cases, the logical formula constituting the search formula may be specified by the user. Furthermore, synonyms or similar words of the important words may be used as keywords. It is preferable that the search formula creation unit 111 saves the created search formula in the search formula DB 122. Furthermore, the important words may be terms or phrases identified from the input text 1000 to be used in the search formula.
[0023] Furthermore, the search execution unit 112 searches for documents using the search query created by the search query creation unit 111. To this end, the search execution unit 112 breaks down the acquired search query into its constituent logical sums, and searches the document DB 121 for each logical sum. The search execution unit 112 saves a document set resulting from the search in the search document DB 123 for each logical sum. Furthermore, the search execution unit 112 also executes searches using search query correction candidates corrected by the search query correction candidate creation unit 114 described below, and searches using corrected search queries identified and extracted from these. At this time, the search execution unit 112 or the search query correction candidate creation unit 114 creates a difference set indicating the difference between the search query before correction and these.
[0024] The search information visualization unit 113 also identifies a set of searched documents for each disjunction, i.e., a search set and a disjunction corresponding to each search set, and creates displayable search information including the disjunction. The search information includes the disjunction and the corresponding search expression. The search information is then displayed, more preferably illustrated.
[0025] Furthermore, the search query correction candidate creation unit 114 generates search query correction candidates, which are correction candidates for the search query, from the search query candidates in the search query DB 122, the search document set in the logical sum search document DB 123, or the correct document in the correct document DB 124. As a result, the search information visualization unit 113 can visualize the search results of the search query correction candidates and the corrected search queries, as well as a set of differences between them. As a result, information including the differences can be presented. The search query creation device 100 may be realized as a search device that searches for search targets such as documents.
[0026] This concludes the explanation of the configurations of the embodiments, and the processing content in embodiment 1 will now be described. Fig. 3 is a flowchart showing the processing procedure in embodiments 1 to 4. Details of each step will be explained while indicating which functional unit in Fig. 1 executes each step, and while indicating the correspondence with the user interface in Fig. 2, the flowcharts in Figs. 4 to 9, and the schematic diagrams in Figs. 10 to 13.
[0027] (1) Search Query Creation Process (Step S31) In the flowchart of Fig. 3, the search query creation unit 111 executes a search query creation process. The details will be described below with reference to Fig. 4. Fig. 4 is a flowchart showing the procedure of the search query creation process in the first to fourth embodiments of the present invention. The search query creation unit 111 performs a morphological analysis on the input text 1000 and divides it into words (step S41).
[0028] Next, the search query creation unit 111 extracts important words from the words that have been morphologically analyzed (step S42). A method based on the frequency of occurrence of words, such as TF-IDF, can be used as a method for extracting important words. In this case, the TF-IDF value of each word in a document is calculated in advance using all or part of the documents stored in the document set DB 121, and a TF-IDF matrix is created. When searching, the TF-IDF matrix is updated using frequency information of the words in the input text 1000 when the input text 1000 is input, and the TF-IDF value of each word in the input text 1000 is calculated. This makes it possible to assign an importance score to each word in the input text 1000. In this specification, a method in which TF-IDF is used to extract important words has been described, but other methods may be used as long as they are methods for extracting important words together with importance scores.
[0029] Next, the search query creation unit 111 collects synonyms for each of the extracted important words (step S43). For synonyms, one or more semantically similar words can be collected using a large-scale language model such as BERT or GPT. At this time, it is desirable to assign a similarity score to each important word indicating how similar the collected synonyms are. The method of collecting synonyms is not limited to this method, and a synonym dictionary prepared in advance may be used. Furthermore, the synonyms may include related words such as so-called synonyms.
[0030] Furthermore, the search formula creation unit 111 associates the key words and their synonyms acquired in steps S42 and S43 with their importance scores and similarity scores, respectively, and stores them in the search formula DB 122 (step S44). Here, Fig. 5 is a diagram showing an example of the search formula DB 122 in the first to fourth embodiments of the present invention.
[0031] The keyword string of the search formula DB 122 stores each keyword and its importance score, and the synonym string stores a list of synonyms collected for a keyword in the same row and their similarity scores. The number of keyword words extracted from the input text 1000 and the number of synonyms collected by the search formula creation unit 111 may be specified by the user or may be set by the system, that is, by each device such as the search formula creation device 100.
[0032] Next, the search formula creation unit 111 creates a search formula from the important words and synonyms stored in the search formula DB 122 (step S45). In this step, the number of important words and synonyms that compose the search formula can be determined arbitrarily. For example, a search formula using the number of important words designated by the user in order of importance score may be created, or the lower limit of the importance score of the important words that compose the search formula may be preset on the system side.
[0033] For example, if the number of key words used in the search formula is three and the number of synonyms is two, the search formula "(car window+window)+(desk+table)+(task+work)" is created using the search formula DB 122. In the search formula creation process in step S31, the results are displayed using, for example, the interface of the search formula creation screen 2000 in FIG. 2. The text input unit 2001 in FIG. 2 accepts text specified by the user, and the search formula creation unit 111 creates the search formula based on this. In addition, the search formula display unit 2002 displays the created search formula in text for each logical sum.
[0034] (2) Search process for each search expression logical sum (step S32) Next, the search execution unit 112 searches the document DB 121 for each logical sum of the search formula created in step S31. For example, for the search formula "(car window + window) x (desk + table) x (task + work)", searches are performed for "car window + window", "desk + table", and "task + work". Then, for the search results, each logical sum and the obtained document are stored in the search document DB 123 for each logical sum.
[0035] (3) Visualization process for each query expression logical sum (step S33) Next, the search information visualization unit 113 obtains the logical sum and documents searched by each logical sum from the logical sum-by-logical sum search document DB 123, and creates search information including a search result set. This search information represents the size of each set of logical sums as an area of a region, and is information indicating the product of the three logical sums, that is, the search result set of the search expression "(car window + window) x (desk + table) x (task + work)" generated in step S31, and the overlap of the three regions. This search information is then visualized information, and is presented to the user.
[0036] An example of visualizing each logical disjunction and the corresponding document set area using circles is shown in the visualization screen 2003 in Fig. 2. The number of search results for the above search formula is 20, and the number of search results for each logical disjunction is 100 for "car window + window," 75 for "desk + table," and 80 for "task + work." A set consisting of a combination of two logical disjunctions is represented by the overlapping area of two circles, and the overlapping area of three circles represents the search result set for the above search formula.
[0037] Generally, a diagram that shows such a search set is called a Venn diagram, but the display method in this specification is not limited to a Venn diagram, and any other method that shows a set of logical sums may be used. Also, the visualization screen 2003 shows logical sums to which keywords have been added as circles surrounded by dotted lines, and this will be explained in detail in the next (4) Search Formula Modification Candidate Presentation Process (step S35).
[0038] (4) Search Query Modification Candidate Creation Process (Step S35) Next, the search query correction candidate creation unit 114 creates a search query correction candidate, which is a correction candidate for the search query, from the search query stored in the search query DB 122 or the correct document stored in the correct document DB 124. The search query correction candidate creation unit 114 also specifies a difference set of the search query correction candidate and the search result documents that change due to a search using the search query correction candidate by adding it to the set created in step S33. It is preferable that the difference set is a result of the search performed by the search execution unit 112 using the search query correction candidate. As a result, the difference set is presented to the user. Here, a method of creating a search query correction candidate from the search query stored in the search query DB 122 will be described, and a method of creating a search query correction candidate from the correct document stored in the correct document DB 124 will be described later.
[0039] 6 is a flowchart showing the procedure of the search expression correction process in the first embodiment of the present invention. First, the search expression correction candidate generating unit 114 acquires a search expression from the search expression DB 122 (step S61). In the acquired search expression, as shown in FIG. 5, each of the key words and synonyms includes an importance score and a similarity score. Next, the search query correction candidate creating unit 114 determines the search query correction parts based on the importance score and similarity score included in the search query (step S62). For this purpose, the search query correction candidate creating unit 114 creates a search query correction policy. There are, for example, the following four search query correction policies:
[0040] [Policy 1] Add keywords in the disjunction. For example, adding the keyword "table" to the disjunction of "desk + table" expands the search set to "desk + table + table."
[0041] [Policy 2] Delete keywords in the disjunction. For example, the keyword "window" is deleted from the disjunction "car window + window + window" to make it "car window + window," thereby reducing the search set.
[0042] [Policy 3] Remove logical sums. For example, from a search expression such as "(car window + window) x (desk + table) x (task + work)," remove "task + work" and make it "(car window + window) x (desk + table). This expands the search set.
[0043] [Policy 4] Add a logical sum. For example, for a search expression such as "(car window + window) x (desk + table) x (task + work)", add "screen + display" to make it "(car window + window) x (desk + table) x (task + work) x (screen + display)". This reduces the search set.
[0044] Then, the search query correction candidate generating unit 114 selects at least one of the correction policies (policy 1 to policy 4) and determines the parts of the search query to be corrected. The policy 1 to policy 4 on which the parts of the search query to be corrected is determined can be determined based on the importance score of the key words, the similarity score of the synonyms, and the number of search documents for the current search query. For example, when the number of search documents is huge and it is necessary to narrow down the search results, [Policy 2] or [Policy 4] is selected. Conversely, when the number of search documents is small and there is a possibility that some documents may have been missed, it is necessary to expand the search set, in which case [Policy 1] or [Policy 3] is selected.
[0045] The process of determining the part of the search query to be revised when the number of retrieved documents is small will be described below. At this time, the revision policy is determined based on whether the importance score of the keyword with the second highest importance after the keyword used in the current search query is less than a threshold, or the similarity score of the synonym with the second highest similarity score after the synonym used in the current search query is equal to or greater than a threshold.
[0046] For example, assume that the current search formula is "(car window + window) x (desk + table) x (task + work)", the importance score threshold is set to 0.7, and the similarity score threshold is set to 0.8. In this case, in the example of Figure 5, the importance scores of the important words "car window", "desk", and "task" used in the current search formula are all above the threshold of 0.7, so the deletion of important words, that is, the deletion of logical sums (policy 3), does not occur.
[0047] In addition, in terms of synonyms, the synonyms with the next highest similarity scores after the synonym used in the current search formula are "window," "table," and "business." Of these, the synonyms that exceed the threshold of 0.8 are "window" and "table," but since "table" has the highest similarity score, a correction is presented to expand the search set by changing "desk+table" to "desk+table+table" (Policy 1). This concludes the explanation of step S62.
[0048] Next, search expression correction candidate creation unit 114 identifies a set to be searched for using the search expression correction candidate "desk+desk+table" and the logical sum "desk+desk+table." To this end, search execution unit 112 executes a search using these search expression correction candidates. Next, search expression correction candidate creation unit 114 accepts the results and identifies the set (step S63). In response to this, the search information visualization unit 113 represents the specified set as an area and creates a search set, which is information that can be illustrated. As a result, the created search set is presented to the user. The search set is included in the above-mentioned search information.
[0049] Furthermore, the search query correction candidate creation unit 114 identifies the search set expanded by correcting the logical sum, that is, the difference set from the current search query (step S64). In response to this, the search information visualization unit 113 creates presentable search information including the difference set, and as a result, the created search information is presented to the user.
[0050] The result of the above search formula correction will be explained with reference to FIG. 12. FIG. 12 is a diagram that shows a schematic representation of the expansion of a search set in the first to fourth embodiments of the invention. In the current search formula (before correction), the search set is region 1203 (20 items) where region 1201, region 1205, and region 1206 overlap. In contrast, region 1201 is expanded to region 1202 by adding "table" to the logical sum "desk+table" that searches region 1201. Furthermore, region 1204 is added to the search set. Moreover, visualization screen 2003 in FIG. 2 is a screen that shows the presenting of search formula correction candidates after this expansion.
[0051] Similarly, even if the search formula correction portion determined in step S62 is different, the difference set is displayed in the same way. For example, as in the above example, consider a case where the search set is small when the current search formula is searched for as "(car window + window) x (desk + table) x (task + work)" and the threshold of the importance score is 0.8. Since the importance score of "work", which has the smallest importance score among the important words used in the current search formula, is 0.71 and is less than the threshold of 0.8, the important word is deleted, that is, the logical sum is deleted (policy 3). In this case, in FIG. 13, the area 1303 surrounded by the dotted line is deleted, and the area 1302 (20 items) representing the search set searched for by the current search formula is expanded, and the area 1301 (25 items) is added as a difference set.
[0052] Similarly, if the search set of the current search expression is large and further narrowing is required, the search set is reduced by [Policy 2] or [Policy 4], but the policy can be determined based on the importance score and similarity score as in the above example. For example, assume that the current search expression is "(car window + window) x (desk + table)". In this case, the search set is the union of area 1302 and area 1301. If the threshold of the importance score is set to 0.7 and all important words above the threshold are added as a logical sum, "work + work" is added as a new logical sum by the search expression shown in Figure 5 (Policy 4). To illustrate this, area 1303 surrounded by a dotted line in Figure 13 is added, and the set searched by the current search expression (the union of area 1302 and area 1301) is narrowed down to only area 1302 by adding the logical sum. In other words, it becomes a difference set from which area 1301 is deleted.
[0053] Similarly, the current search formula is "(car window + window) x (desk + table + table) + (task + work)", and the threshold value of the similarity score is set to 0.9, and synonyms less than the threshold value are deleted from the logical sum. As a result, "table" is deleted from "desk + table + table" by the search formula shown in FIG. 5. To illustrate this, in FIG. 12, area 1202 searched for with "desk + table + table" is reduced to area 1201 searched for with "desk + table", and the search set is also reduced from the union of area 1203 and area 1204 to only area 1203. In other words, area 1204 becomes the difference set from which it is deleted. This concludes the explanation of step S35.
[0054] (5) Query Feedback Processing (Step S36) Next, the feedback unit 115 determines whether to modify the search query, that is, which of the search query modification candidates to use as the modified search query, based on the feedback on the current search set (search results using the search query before modification) by the search query modification candidate generating unit 114. This feedback includes the user's designation of documents in the difference set to be added to or deleted from the current search set. The procedure for this search query feedback process will be described below with reference to the flowcharts in Figs. 7, 8, and 9.
[0055] First, the flowchart in Fig. 7 will be described. The search query feedback process is started by a user operation. For example, the user checks the search query correction candidates presented on the visualization screen 2003 in Fig. 2, clicks on the difference set, and clicks on the "Feedback" button to start the feedback process.
[0056] First, the feedback unit 115 acquires documents in the difference set (step S71). At this time, it is determined whether the number of acquired documents is equal to or greater than a predetermined threshold (N documents) (step S72). As a result, if the number is equal to or greater than N (YES), the process proceeds to step S73. If the number is less than N (NO), the process proceeds to step S74.
[0057] Then, the feedback unit 115 samples the acquired documents so that the number is less than N (step S73). This is because if there are a large number of documents to provide feedback on, the burden on the user increases. Therefore, in step S73, documents are sampled from the difference set, but the sampled document set needs to be at least in number and be a document set sampled without bias from the difference set. This is to provide accurate search query feedback even with a small number of documents sampled.
[0058] To perform unbiased sampling from the difference set, it is necessary to sample documents that contain as many different combinations of search keywords as possible. For example, consider the current search formula being (car window + window) x (desk + table) x (task + work), and sampling from the difference set when a new keyword "table" is added to the logical OR "desk + table." In this case, in addition to "table," the documents in the difference set should always contain both "car window" or "window," and "task" or "work," meaning there are four possible keyword combinations. Figure 10 shows an example list of keyword combinations for documents in this difference set.
[0059] In this case, document numbers 021 and 064 contain the same search keyword combination, "view from train window x work," so they are likely to be documents with similar content, and it would be effective to sample one of them.Similarly, document numbers 135, 175, 205, and 208 all contain the search keyword combination, "window x work," so any number of documents with similar overlapping combinations can be sampled from among them.
[0060] Suppose we sample one document for each of the four keyword combinations, and randomly sample documents with the same keyword combinations. In this way, for example, from the combination list in Figure 10, we can sample four documents with document numbers 021, 092, 105, and 208.
[0061] The sampling method is not limited to this, and other methods may be used, such as sampling a number proportional to the number of documents corresponding to each keyword combination (for combinations with a large number of documents, sampling is performed in proportion to the ratio). Furthermore, the sampling method may be other than the above-mentioned method. Also, steps S72 and S73 may be omitted.
[0062] Next, the feedback unit 115 accepts the user's evaluation of the sampled document or the acquired document as to whether it is the desired document (step S74). An example of the interface used when accepting this evaluation is shown in FIG. 11. The text that the user refers to when evaluating may be the entire target text, or for each document, the first sentence of the document, a representative sentence, or a summary prepared in advance may be displayed. In response, the user selects ◯ if the document is the desired document, or × if not, to provide feedback of correctness or incorrectness.
[0063] Next, the feedback unit 115 tally up the evaluation results (step S75). To do this, the feedback unit 115 counts the number of documents labeled as correct and incorrect. From this point on, this tallying up result is used to decide whether to accept the proposed correction candidate search formula (actually correct the search formula) or not (keep the current search formula as is). However, the process flow differs depending on whether the correction candidate search formula expands or shrinks the search set.
[0064] Therefore, the feedback unit 115 judges whether to expand the search result set (step S76). If the search result set is to be expanded (YES), (1) the process transitions to the process of the flowchart in Fig. 8. More specifically, if the search query correction candidate creation unit 114 suggests a search query correction in [Strategy 1] or [Strategy 3], the process transitions to the flowchart in Fig. 8.
[0065] If the search result set is to be reduced (NO), (2) the process transitions to the flowchart in Fig. 9. More specifically, if the search query correction candidate creation unit 114 presents a search query correction in [Strategy 2] or [Strategy 4], the process transitions to the flowchart in Fig. 9.
[0066] 8 and 9 will be described below. First, the processing procedure in the flowchart in Fig. 8 will be described. The feedback unit 115 judges whether the number of correct documents counted in step S75 is equal to or greater than a predetermined threshold M (step S81). As a result, if the number is equal to or greater than M (YES), the process proceeds to step S84. If the number is less than M (NO), the process proceeds to step S82.
[0067] Furthermore, the feedback unit 115 judges whether the expanded difference set contains one or more correct answer documents (step S82). As a result, if one or more are included (YES), the process proceeds to step S83. If not, the process ends. This is because if even one correct answer document is included, the correct answer document will be missed, so the correct answer document is stored in the correct answer document DB and the search set is expanded by another method in the process described later.
[0068] Furthermore, the feedback unit 115 stores the correct document in the correct document DB (step S83) and ends the feedback process. In this case, the correction of the search formula is omitted. This is because the set expanded by the correction of the search formula contains many noise documents, and the search accuracy decreases if the correction of the search formula is accepted.
[0069] The feedback unit 115 also modifies the search formula. To this end, the feedback unit 115 expands the search set by adding keywords in the disjunction or deleting the disjunction (step S84). This is because the set to be expanded will contain many correct documents, and the search accuracy can be improved by accepting the modification of the search formula and expanding the search set.
[0070] Next, the flowchart in Fig. 9 will be described. First, similarly to step S81, the feedback unit 115 judges whether the number of correct answer documents counted in step S75 is equal to or greater than the predetermined threshold M (step S91). As a result, if there are M or more (YES), the process ends without modifying the search formula. This is because if the reduced search set contains many correct answer documents, the reduction will result in a decrease in search accuracy. On the other hand, if the number of correct answer documents is less than M (NO), the process proceeds to step S92.
[0071] The feedback unit 115 also determines whether the difference set contains one or more correct answer documents (step S92). As a result, if one or more correct answer documents are included (YES), the process proceeds to step S93. If not, the process proceeds to step S94. The reason for this process is that If even one correct answer document is included, the correct answer document will have been omitted, so step S93 described below is executed and the search set is expanded in a different manner in the process described below, similar to FIG.
[0072] The feedback unit 115 also stores the supervised document in the supervised document DB 124 (step S93). Then, the feedback unit 115 corrects the search formula in the same way as in step S84. To this end, the feedback unit 115 reduces the search set by deleting keywords in the logical sum or adding logical sums (step S94). This is because the reduced search set contains a lot of noise, and thus reducing it can improve search accuracy.
[0073] (6) Document list output process (step S37) Finally, the document list output unit 116 outputs a list of documents searched using the search query whose correction has been confirmed by the feedback unit 115. This output includes display on the search query creation device 100 and transmission to another device. In the former case, the document set list is displayed on a display unit such as a display included in the search query creation device 100 or a display unit directly connected thereto. In the latter case, the document set list is transmitted to another device such as a terminal device. As a result, it can be displayed on another device.
[0074] Here, it is assumed that the processes (1) to (5) are repeated. That is, the feedback unit 115 receives feedback on whether or not to accept the correction candidates created by the search query correction candidate creation unit 114 and presented to the user. After receiving the feedback, the visualization screen 2003 in FIG. 2 is updated, and the next search query correction candidate is presented. This is continued until the user selects to end the search query correction (step S34). In addition, the display in the document list output process in (6) may be performed by the user selecting "display document list" on the visualization screen 2003 while the user is correcting the search query, or may be performed when the user ends the search query correction.
[0075] In the above explanation of (4), only the processing method in which the search query correction candidate generating unit 114 presents search query correction candidates from the search queries stored in the search query DB 122 was described, but a process of generating search query correction candidates from the correct documents stored in the correct document DB 124 may also occur. This occurs when, in the above-mentioned feedback process, a small number of correct documents are included in the difference set, but the number of correct documents is below a predetermined threshold, so the search query is not corrected to add the difference set to the search set. That is, this is an example in the flowchart of FIG. 8 where the number of correct documents is less than the threshold, and the search query correction is not accepted and the search set is not expanded, or an example in the flowchart of FIG. 9 where the number of correct documents is less than the threshold, and the search query correction is accepted and the search set is reduced.
[0076] In these cases, after the processing by the feedback unit 115, the search query correction candidate generating unit 114 presents a correction to the search query so that the search result set includes the correct document, in accordance with the processing procedure in the flowchart of Fig. 14. This processing will be described below.
[0077] First, the search query correction candidate creation unit 114 acquires a correct answer document from the correct answer document DB 124 (step S141). Next, the search query correction candidate creation unit 114 counts the DF value (the number of documents in which the word appears) for each word included in the correct answer document other than the word included in the current search query (step S142). Furthermore, the search query correction candidate creation unit 114 acquires the word with the maximum DF value in each document (step S143). Also, if there is a word that commonly appears in the correct answer document set, the word with the maximum DF value in each document tends to overlap. Therefore, the search query correction candidate creation unit 114 creates a logical sum from the words collected in this way (step S144). For example, if the words with the maximum DF value in each document in the correct answer document set are "screen", "display", and "video", "screen + display + video" is created.
[0078] Next, the search query correction candidate creation unit 114 judges whether the addition of a logical sum was presented in the preprocessing (correction candidate presentation process). Here, the presentation of the addition of a logical sum means [Policy 4]. In this case, the answer is YES in the figure, and the process transitions to step S146. Also, if the preprocessing presented the deletion of a keyword from the logical sum [Policy 2], the answer is NO in the figure, and the process transitions to step S147.
[0079] Furthermore, the search formula correction candidate creating unit 114 adds the generated logical sum to the logical sum added in the preprocessing (step S146). For example, if the logical sum added in the preprocessing was "task+work", then "screen+display+video" is added to it to make it "task+work+screen+display+video". Furthermore, the search formula correction candidate creating unit 114 adds the generated logical sum to the logical sum from which the keyword was deleted in the preprocessing (step S147). For example, if "business" was deleted from "task+work+business" in the preprocessing to make it "task+work", then the logical sum "screen+display+video" created in step S144 is added to make it "task+work+screen+display+video".
[0080] With this process, if the difference set contains a correct answer document but has been reduced or not expanded, a new logical sum is created using the keywords in the correct answer document contained in the difference set, and this is added as a sum to the current search query. This expands the search set so that the correct answer document is not missed.
[0081] As described above, according to the first embodiment, the search formula generated from the input text 1000 and its set are illustrated, along with the difference set that is changed by correcting the search formula correction candidates. Then, by providing feedback to a small number of documents in the difference set, the determination of whether or not the search formula can be corrected is semi-automated, making it easy to brush up the search formula. EXAMPLES
[0082] Next, in the second embodiment, the search set is reduced in the process of presenting search query correction candidates. To reduce the search set, keywords in the OR search search query are deleted, logical sums are added, or logical sums are multiplied. More specifically, the process of reducing the search set in the first embodiment ([Policy 2] and [Policy 4] in the first embodiment) is performed based on the user's feedback results for the search set of the current search query. The details are described below.
[0083] 15 and 16 are flowcharts showing the procedure of the search query correction candidate generation process in this embodiment. These flowcharts will be explained below, but since the process is the same as that in the flowchart in FIG. 3 except for step S35, it will be omitted.
[0084] First, we will explain the processing of the search query correction candidate creation unit 114 in the flowchart shown in Fig. 15. More specifically, this flowchart in Fig. 15 shows a processing procedure for presenting whether the search result set should be reduced, and if so, which keyword in which logical sum in the current search query should be deleted to reduce the search result set created by that logical sum.
[0085] First, the search query correction candidate generating unit 114 acquires documents of the current search set (step S151). If the current search query is "(car window + window) x (desk + table) x (task + work)", the current search set corresponds to the area 1206 in FIG. 12, which contains 20 documents. If the number of documents acquired here is equal to or greater than a threshold, document sampling may be performed in step S73 by the feedback unit in the first embodiment. Here, a description of the sampling process in step S73 will be omitted.
[0086] Next, the search query correction candidate creating unit 114 accepts the user's evaluation of the documents (step S152). The user evaluates whether each document is correct or incorrect in the example interface of FIG. 11, similar to the process performed by the feedback unit in the first embodiment. The search query correction candidate creating unit 114 also acquires each of the correct and incorrect documents from the evaluation results (step S153). The search query correction candidate creating unit 114 then acquires the search keywords contained in each document (step S154). Here, the search keywords refer to the words contained in each document that are contained in the current search query (i.e., the words that matched in the search).
[0087] Next, the search query correction candidate creation unit 114 calculates the DF value of the search keyword for each of the correct document set and the incorrect document set, and obtains a search keyword whose DF value is equal to or greater than a threshold value in the incorrect document set and whose DF value is 0 in the correct document set (step S155). In other words, it obtains a search keyword that frequently appears in the incorrect document set but does not appear in the correct document set. Then, the search query correction candidate creation unit 114 decides to delete the search keyword from the keywords in the logical sum of the current search query. This result is presented to the user (step S156). Note that the presentation of the keyword deletion is accompanied by a display of the difference between the screen and the search set, as in the first embodiment, but since this is the same process, a description thereof will be omitted here.
[0088] An example of the process in steps S155 and S156 is shown below. If the search keyword acquired in step S155 is "work", then "work" is a keyword that does not appear in documents that the user has determined to be correct, but appears a certain number of times in documents that the user has determined to be incorrect. Therefore, in step S156, "work" is deleted from the search logical sum "work + work" that includes "work". As a result, a suggestion is made to modify the search formula in a direction that narrows the search set and reduces noise.
[0089] Next, a description will be given of Fig. 16. This flowchart shows a processing procedure for determining whether the search result set should be reduced, and if so, presenting the logical sum to be added (multiplied as a product) to the current search expression.
[0090] First, the processing procedure from step S161 to S163 performed by the retrieval query correction candidate creation unit 114 is the same as steps S151 to S153 in the flowchart in FIG. 15, so a description thereof will be omitted. Next, the retrieval query correction candidate creation unit 114 acquires words other than the search keywords for each of the acquired correct answer document set and incorrect answer document set (step S164). Next, the retrieval query correction candidate creation unit 114 calculates the DF value of the word for each of the correct answer document set and the incorrect answer document set, and acquires the search keywords whose DF is equal to the number of correct answer documents in the correct answer document set and whose DF is equal to or less than a threshold value in the incorrect answer document set (step S165). In other words, the words that appear in all documents in the correct answer documents and that appear only at a frequency less than the threshold value in the incorrect answer document set are acquired as search keywords to be added as new conditions.
[0091] Next, the search query correction candidate creation unit 114 acquires synonyms of the keywords (step S166). To this end, it executes a process similar to that of the search query creation unit 111 in the first embodiment (step S43 in the flow chart of FIG. 4). In addition, the search query correction candidate creation unit 114 generates a logical sum. To this end, it executes a process similar to that of the search query creation unit 111 in the first embodiment (step S45 in the flow chart of FIG. 4) from the key words and synonyms.
[0092] The search query correction candidate creation unit 114 also adds the logical sum to the current search query (step S167). In other words, the search result set is reduced by overlapping a set consisting of the new logical sum. The addition of the logical sum (reduced search result set) is then presented. The presentation of the addition of the logical sum is accompanied by a display of the difference between the screen and the search result set, as in the first embodiment, but as this is the same process, a description thereof will be omitted here.
[0093] An example of the process from step S165 to step S167 is shown below. As in the explanation of FIG. 15, assume that the current search formula is "(car window + window) x (desk + table) x (task + work)". If the word acquired in step S165 is "monitor", "monitor" is a word that appears commonly in all documents that the user judged to be correct, and appears less than a certain number of times in documents that the user judged to be incorrect. Therefore, in step S166, synonyms are created from "monitor" to obtain "screen" and "display". Then, in step S167, a logical sum "monitor + screen + display" is newly added, and it is proposed to modify the search formula in a direction that narrows the search set and reduces noise.
[0094] As described above, according to the second embodiment, the correction candidates for the current search query can be specified based on the user's feedback results for the set obtained by the current search query. Moreover, by presenting the correction candidates, it is possible to present a more accurate correction for the search query compared to the first embodiment. EXAMPLES
[0095] In the third embodiment, a vector search is used to change the search formula. At this time, a search logical sum is added or deleted in the process of presenting search formula correction candidates in the first embodiment. The search information visualization unit 113 creates a process for reducing the search result set [strategy 4] or expanding it [strategy 3], and performs a vector search on each of these search result sets.
[0096] 17 and 18 are flowcharts showing the procedure of the search query correction candidate generation process in this embodiment. Note that, in these flowcharts, the processes other than step S35 are the same as those in the flowchart in FIG. 3, so descriptions of these processes will be omitted.
[0097] First, a description will be given of the flowchart shown in Fig. 17. This flowchart shows a processing procedure for presenting whether the search result set should be reduced, and if so, the logical sum to be added (multiplied as a product) to the current search expression.
[0098] First, the search query correction candidate generating unit 114 acquires documents of the current search result set (step S1701). If the current search query is "(car window+window)×(desk+table)×(task+work)", the current search result set corresponds to the area 1206 in Fig. 12, which contains 20 documents. In this embodiment, the number of documents acquired here does not need to be considered because the user does not intervene.
[0099] Next, the search query correction candidate creation unit 114 acquires vectors for each document (step S1702). The acquisition, or conversion, of vectors for each document can be achieved using existing vectorization techniques such as TF-IDF and BERT. The search query correction candidate creation unit 114 also vectorizes the input text 1000 to obtain input text vectors (step S1703). Next, the search query correction candidate creation unit 114 measures (calculates) the similarity of each document vector to the input text vector. To achieve this, the search query correction candidate creation unit 114 uses existing inter-vector similarity measurement indices such as cosine similarity (step S1704).
[0100] Next, the search query correction candidate generating unit 114 classifies each document according to its similarity. For example, documents with similarity equal to or greater than threshold A and documents with similarity less than threshold B are obtained, and designated as documents A and B, respectively (step S1705). Each threshold may be determined based on the similarity measurement index used. For example, in the case of cosine similarity, the value is between -1 and 1, and the closer to 1, the more similar. For this reason, by setting threshold A to a value close to 1, such as 0.8, and threshold B to a value close to -1, such as -0.5, it is possible to obtain only documents similar to the input text as documents A, and only documents not similar to the input text as documents B.
[0101] Next, the search formula correction candidate creation unit 114 acquires words other than the search keyword from the document acquired in step S1705 (step S1706). Next, the search formula correction candidate creation unit 114 calculates the DF value of each word for each of document A and document B, and acquires words whose DF is equal to or greater than threshold C in document A and equal to or less than threshold D in document B (step S1707). For example, if the number of documents A and B is 20, the threshold C is 15, and the threshold D is 5, words that appear in 15 or more documents in document A and appear in 5 or less documents in document B are acquired. Note that thresholds C and D are not limited to the number of documents, and may be percentages, and can be set to any value.
[0102] The search query correction candidate creation unit 114 also uses this word as a key word and collects synonyms in the same manner as the search query generation unit in the first embodiment (step S1708). The search query correction candidate creation unit 114 then generates a logical sum from the key word and synonyms in the same manner as the search query generation unit in the first embodiment, and adds the logical sum to the current search query. In this way, the search result set is narrowed (reduced) by superimposing a set consisting of the new logical sum (step S1709). As a result, the reduced search result set is presented. The presentation of the logical sum addition is accompanied by a display of the difference between the screen and the search result set, as in the first embodiment, but as this is the same process, a description thereof will be omitted here.
[0103] An example of the process from step S1707 to step S1709 is shown below. A case will be described where the current search formula is "(car window + window) x (desk + table) x (task + work)" and the word obtained in step S1707 is "video". In this case, "video" is a word that appears a certain number of times or more in documents similar to the input text and a certain number of times or less in documents not similar to the input text. Therefore, in step S1708, synonyms are created from "video" to obtain "image" and "map". Then, in step S167, the logical sum "video + image + map" is newly added, and it is determined that the search formula should be modified in a direction that narrows the search set and reduces noise, and this is presented.
[0104] Next, a description will be given of the flowchart shown in Fig. 18. This flowchart shows a processing procedure for presenting whether the result set should be expanded and, if so, which search disjunctions should be deleted to expand the result set.
[0105] First, the search formula correction candidate generating unit 114 acquires documents in a difference set when one logical sum is removed from the current search formula (step S181). For example, if the current search formula is "(car window + window) x (desk + table) x (task + work)", the current search set corresponds to area 1206 in Fig. 12, and there are three difference sets. One is the intersection of areas 1201 and 1205 minus area 1206, another is the intersection of areas 1204 and 1205 minus area 1206, and yet another is the intersection of areas 1204 and 1201 minus area 1206.
[0106] Next, search query correction candidate preparation unit 114 performs the same processes as steps S1702 and S1703 (steps S182 and S183). As a result, the vectors of each document and the vector of input text 1000 are calculated. Then, search query correction candidate preparation unit 114 calculates the similarity of each document vector to the input text vector, as in step S1704 (step S184).
[0107] Next, the search query correction candidate creation unit 114 calculates the average of the similarity obtained in step S184 for each difference set (step S185). Then, the search query correction candidate creation unit 114 identifies, as a deletion candidate logical disjunction, a logical disjunction to be deleted in order to expand the difference set whose average similarity is the largest and is equal to or greater than the threshold value. This deletion logical disjunction is presented (step S186). As in the first embodiment, the identification and presentation of the logical disjunction to be deleted is performed in conjunction with the display of the difference between the screen and the search set, and therefore a description thereof will be omitted here. Also, as in the case of the above, the threshold value of the similarity can be set to a value close to 1, such as 0.6 in the case of cosine similarity.
[0108] An example of the process in step S186 is shown below. Suppose the difference set with the maximum average similarity and above the threshold in step S186 is the intersection of areas 1201 and 1205 in FIG. 12 minus area 1206. This indicates that this difference set is similar to the input text and is highly likely to contain a large set of correct documents, and it is proposed to include this difference set in the search set. In this case, to expand this area, it is necessary to delete the condition in area 1204, that is, to delete the logical sum "task + work". For this reason, it is proposed to delete "task + work" and modify the search formula in a direction that narrows the search set and reduces noise.
[0109] As described above, according to the third embodiment, correction candidates for the current search formula are created by the search information visualization unit 113, and are implemented based on the results of vector search performed on each displayed set. This makes it possible to realize highly accurate search formula correction and presentation thereof while reducing the burden on the user. EXAMPLES
[0110] In this embodiment, a method of using a neighborhood search condition in the process of presenting search query correction candidates in the embodiment 2 will be described. Fig. 19 is a flowchart showing the processing procedure of the process of creating search query correction candidates in this embodiment. In this flowchart, the processes other than step S35 in the flowchart of the embodiment 1 shown in Fig. 3 are the same, so they will be omitted.
[0111] First, the search query correction candidate generating unit 114 executes the same processes as steps S151 and S152 in the second embodiment (steps S191 and S192).
[0112] Next, the search query correction candidate creating unit 114 acquires the correct document based on the user evaluation obtained in step S192 (step S193). At this time, it is desirable to acquire a limited number of correct documents. Next, the search query correction candidate creating unit 114 acquires the search keyword and the appearance position in the correct document (step S194). Here, the appearance position may be a number indicating the number of characters or the number of words counting from the beginning. Next, the search query correction candidate creating unit 114 measures the maximum value of the distance between the search keywords of each logical sum included in the current search query (step S195). Then, the search query correction candidate creating unit 114 identifies the maximum value of the distance obtained in step S195 and the combination of the search logical sum having the maximum value (step S196). Then, this combination is presented.
[0113] An example of the process from step S193 to step S196 is shown below. It is assumed that the current search query is "(car window+window)×(desk+table)×(task+work)" and the following three correct answer documents are acquired in step S193.
[0114] Correct answer: Document 1: Images can be displayed on the train window and work can be done at a desk.
[0115] Correct answer: Document 2: By installing a folding desk behind the seat and a screen display device on the window, it becomes possible to work inside the car.
[0116] Correct answer: Document 3: We provide vehicles that enable desk work while traveling and can project images onto the windows.
[0117] In this case, the search keywords and their appearance positions obtained in step S194 are as follows: Correct document 1: "Car window" is the 0th character, "Desk" is the 12th character, and "Work" is the 19th character. Correct document 2: "Window" is the 16th character, "desk" is the 11th character, and "work" is the 46th character. Correct document 3: "Car window" is the 16th character, "desk" is the 5th character, and "work" is the 7th character. Next, in step S195, the maximum distance between the search keywords of each logical sum is measured. Maximum distance between "car window" and "desk": 8 characters Maximum distance between "desk" and "work": 32 characters The maximum distance between "work" and "car window" is 25 characters.
[0118] Therefore, the neighborhood search formula presented in step S196 is "{(car window + window) x (desk + table), 8n} x {(desk + table) x (task + work), 32n} x {(work + work) x (car window + window), 25n}." This makes it possible to further narrow down the search set by neighborhood search.
[0119] As described above, according to the fourth embodiment, it is possible to present correction candidates for the current search query based on the user's feedback results for the set obtained by the current search query, taking into consideration even the distance of the keywords in the search query. Therefore, it is possible to present corrections for the search query with higher accuracy than in the second embodiment. EXAMPLES
[0120] The fifth embodiment is an implementation example of a document search system in which the search query creating device 100 of the first to fourth embodiments is implemented in a server (computer). Fig. 20 is a system configuration diagram showing an implementation example of the document search system in the fifth embodiment.
[0121] 20, in the document search system, a search query creation device 100 is connected to each site (200 to 400) via a network 500. Note that the number of each device shown in FIG. 20 is an example and is not limited to these.
[0122] First, each site will be described. First, the site 200 is provided with a plurality of terminal devices 201-1 to 201-3. Here, the terminal devices 201-1 to 201-3 can be realized by information processing devices such as PCs, tablets, and smartphones. These devices issue instructions to the search formula creation device 100 and output the processing results of the search formula creation device 100 according to user operations. In other words, these devices have the input function and output function (presentation function) of the first to fourth embodiments. The site 200 may be installed in the same area as the search formula creation device 100, or the search formula creation device 100 and the terminal devices 201-1 to 201-3 may be directly connected without going through the network 500.
[0123] Moreover, the site 300 is a site that realizes a database system, and is provided with a management device 301, a document management server 302, and a storage device 303. It is used to manage the operation of the database system. The document management server 302 has a function of managing the storage device 303 according to instructions from the management device 301 or a preset algorithm. The storage device 303 also stores a document DB 121. However, it may also store other DBs, such as a search query DB 122 to a correct answer document DB 124. However, since each of these DBs may be stored in the search query creation device 100, it may be stored in either the database system of the site 300 or the search query creation device 100. In this case, storage in the other may be omitted.
[0124] Next, the site 400 is also a site that realizes a database system, and is provided with a document management server 401 and a storage device 402. This configuration is the same as the site 300 except for the management device 301. However, at least one of the sites 200 and 400 can be omitted.
[0125] Finally, the search formula creation device 100 will be described. The search formula creation device 100 can be realized by a server in a cloud system or the like. For this purpose, the search formula creation device 100 uses a search formula creation program 141 to realize the functions of each unit shown in Fig. 1. The search formula creation device 100 also includes a communication unit 11, a processing unit 12, a memory 13, and a storage unit 14, which are connected to each other via a communication path.
[0126] First, communication unit 11 transmits and receives information to and from each site via network 500. For example, communication unit 11 transmits the above-mentioned presentation contents (difference sets, etc.) to terminal devices 201-1 to 201-3, and receives documents searched for from document DB 121 of storage device 303. The received information is sent to processing unit 12 and memory 13.
[0127] The processing unit 12 can be realized by a processor such as a CPU (Central Processing Unit), and performs calculations according to various programs expanded in the memory 13 described later. That is, the processing unit 12 performs processing of the search query creation unit 111, the search execution unit 112, the search information visualization unit 113, the search query correction candidate creation unit 114, the feedback unit 115, and the document list output unit 116 according to the search query creation program 141. For this reason, the search query creation program 141 has a search query creation module 142, a search execution module 143, a search information visualization module 144, a search query creation / search query correction candidate creation module 145, a feedback module 146, and a document list output module 147. Note that each of these modules may be realized by an individual program, or may be realized as a function of another program such as a document management program. Note that the document list output unit 116 may be realized by the communication unit 11, and the document list output module 147 may be omitted.
[0128] The storage unit 14 can be realized by a storage such as a hard disk drive, etc. The storage unit 14 corresponds to the storage unit 120 in FIG.
[0129] The search formula creation device 100 may further include an input unit that accepts input from a user and an output unit that outputs various information. The search formula creation program 141 may be distributed to the search formula creation device 100 via the network 500, or may be stored in a storage medium and installed in the search formula creation device 100 from the storage medium.
[0130] This concludes the description of the embodiments of the present invention. The present invention is not limited to these embodiments. For example, the document processing device 10 may be realized as a document search device, or may be realized as a so-called stand-alone device. [Explanation of symbols]
[0131] 100: Search expression creation device 110: Processing section 111: Search Query Creation Department 112: Search execution unit 113: Search information visualization unit 114: Search expression correction candidate creation section 115: Feedback section 116: Document list output section 120: Storage section 121: Document DB 122: Search expression DB 123: Logical sum search document DB 124: Correct document DB 1000: Input text
Claims
1. A search query creation device that supports the creation of a search query for searching documents based on an input text, comprising: a search query creation unit that creates a search query using keywords extracted from the input text; a search execution unit that searches documents using the search query to create a first document set; a search query correction candidate creation unit that creates a search query correction candidate by modifying the search query in accordance with a result of the evaluation of the first document set; A search query creation device having a difference set showing the difference between the first document set and a second document set created using the search query correction candidates, the search query creation device having a feedback unit that receives feedback being a confirmation result of the difference set presented to a user of the search query creation device, and identifies a corrected search query from the search query correction candidates in accordance with the confirmation result.
2. 2. The search query creation device according to claim 1, The search query creation device further comprises a search information visualization unit 113 that creates search information including the search query, the first document set for each search disjunction, the search query correction candidates, and the second document set for each search disjunction.
3. 3. The search query creation device according to claim 2, The search query correction candidate creation unit is a search query creation device that calculates the evaluation result in response to the feedback from a user regarding the difference set.
4. 2. The search query creation device according to claim 1, the search execution unit executes a vector search to generate the first document set, the second document set, and the difference set; The search query correction candidate creation unit is a search query creation device that calculates the evaluation result according to a similarity to the input text contained in the difference set.
5. A search query creation method using a search query creation device that supports the creation of a search query for searching documents based on an input text, comprising: A search query creation unit creates a search query using the keywords extracted from the input text; a search execution unit searches for documents using the search query to generate a first document set; a search query correction candidate creation unit creates a search query correction candidate by modifying the search query in accordance with the evaluation result for the first document set; A search query creation method in which a feedback unit receives feedback that is a confirmation result of the difference set presented to a user of the search query creation device, the difference set indicating the difference between the first document set and a second document set created using the search query correction candidate, and identifies a corrected search query from the search query correction candidate in accordance with the confirmation result.
6. 6. The method for creating a search query according to claim 5, A search query creation method in which a search information visualization unit creates search information including the search query, the first document set for each search logical disjunction, the search query correction candidates, and the second document set for each search logical disjunction.
7. 7. The method for creating a search query according to claim 6, the search query correction candidate creation unit calculates the evaluation result in response to the feedback from a user regarding the difference set.
8. 6. The method for creating a search query according to claim 5, a vector search is executed by the search execution unit to generate the first document set, the second document set, and the difference set; The search query creation method, wherein the search query correction candidate creation unit calculates the evaluation result according to a similarity to the input text contained in the difference set.
Citation Information
Patent Citations
Document information evaluation device, document information evaluation method, and document information evaluation program
WO2021245814A1