Refine search requests for content providers
Through topic model analysis and correlation scores, search requests are automatically refined, and the fuzziness of text-based search query is solved, improving search efficiency and accuracy.
Patent Information
- Application Number
- CN202180034989.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-08
- Filing Date
- 2021-05-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-05-19
AI Technical Summary
In the prior art, text-based search queries are ambiguity, resulting in users requiring multiple iterations to refine the search queries to ensure relevant content is obtained, and expert knowledge is required to efficiently set the logical items of keywords and data sources.
By receiving the initial search request, the topic model is applied to analyze the sample document collection, generate topic representations, and automatically refine the search requests to increase document selectivity for high-relevant topics based on the topic relevance score and source relevance score.
Reduces the number of search iterations, improves user work efficiency, reduces the complexity of building related search results and the possibility of errors excluding unrelated results.
Smart Images

Figure CN115605857B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to searching for computer-based text information, and more particularly, to automatically generating suggestions for refined search requests based on an initial search request. Background Art
[0002] Online media analysis tools obtain documents of a specific data source from content providers such as Twitter, Facebook or SocialGist. Content provider APIs typically support at least two access mechanisms that can often be combined: keyword-based retrieval, where the user specifies a set of keywords, optionally combined with AND, OR or NOT, and the data provider returns documents containing the content, such as websites, text documents, forum posts, blog entries, etc.; and site-based retrieval, where the user specifies a list of data sources such as websites, website sections, channels, feeds, etc., and the content provider returns documents specifically from these sources.
[0003] In response to an input search request, document samples are typically presented to the user. Before running a full analysis, the user can continue to search for a more relevant set of documents that better support the target analysis. In particular, when searching for user-generated content, keyword searches may result in ambiguous search results because a large amount of content can be found in social media data sources. For example, a search for "F-50" can return content on specific models of sports cars, soccer cleats, turboprop aircraft, and coffee machines. To this end, the user can refine the query by adding keywords and / or sites. Summary of the Invention
[0004] In one aspect, a computer-implemented method for refining an initial search request to a content provider is disclosed. The computer-implemented method includes receiving an initial search request from a user and submitting the initial search request to a content provider. The computer-implemented method further includes receiving a set of sample documents and source identifiers of corresponding sample documents in the sample documents from the content provider, where the source identifiers identify corresponding data sources in the data source associated with the corresponding sample documents in the sample documents. The computer-implemented method further includes applying a topic model to the set of sample documents to obtain a topic representation, where the topic representation is a description of the topics covered by the corresponding sample documents in the sample documents. The computer-implemented method further includes presenting the topic representation to the user and receiving topic relevance scores of corresponding topics in the topics from the user. The computer-implemented method further includes classifying the data sources according to the topic relevance scores to obtain source relevance scores of corresponding data sources in the data source. The computer-implemented method further includes determining a refined search request based on the source relevance scores, the refined search request having increased selectivity for documents covering the topic with the highest score in the topics.
[0005] In another aspect, a computer program product for refining an initial search request to a content provider is disclosed. The computer program product includes a computer-readable storage medium having program instructions embodied therein, and the program instructions are executable by one or more processors. The program instructions are executable to receive an initial search request from a user. The program instructions are further executable to submit the initial search request to a content provider. The program instructions are further executable to receive, from the content provider, a collection of sample documents and source identifiers of corresponding sample documents in the sample documents, where the source identifiers identify corresponding data sources in a data source that are associated with the corresponding sample documents in the sample documents. The program instructions are further executable to apply a topic model to the collection of sample documents to obtain a topic representation, where the topic representation is a description of topics covered by the corresponding sample documents in the sample documents. The program instructions are further executable to present the topic representation to the user. The program instructions are further executable to receive a topic relevance score of a corresponding topic in the topic from the user. The program instructions are further executable to classify the data sources according to the topic relevance score to obtain a source relevance score of the corresponding data sources in the data source. The program instructions are further executable to determine, based on the source relevance score, a refined search request that has increased selectivity for documents covering the topic with the highest score in the topic.
[0006] In yet another aspect, a computer system for refining an initial search request to a content provider is disclosed. The computer system includes one or more processors, one or more computer-readable tangible storage devices, and program instructions stored on at least one of the one or more computer-readable tangible storage devices for execution by at least one of the one or more processors. The program instructions are executable to receive an initial search request from a user; submit the initial search request to a content provider; receive, from the content provider, a collection of sample documents and source identifiers of corresponding sample documents in the sample documents, the source identifiers identifying corresponding data sources in a data source that are associated with the corresponding sample documents in the sample documents; apply a topic model to the collection of sample documents to obtain a topic representation, the topic representation being a description of topics covered by the corresponding sample documents in the sample documents; present the topic representation to the user; receive a topic relevance score of a corresponding topic in the topic from the user; classify the data sources according to the topic relevance score to obtain a source relevance score of the corresponding data sources in the data source; and determine, based on the source relevance score, a refined search request that has increased selectivity for documents covering the topic with the highest score in the topic. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1Shows an exemplary computing environment adapted to retrieve sample documents in response to a search request according to an embodiment of the present invention.
[0008] Figure 2 Shows an exemplary topic representation according to an embodiment of the present invention.
[0009] Figure 3 Is a diagram showing the processing of relevance scores according to an embodiment of the present invention.
[0010] Figure 4 Is a flowchart showing the steps of a method for refining an initial search request to a content provider according to an embodiment of the present invention.
[0011] Figure 5 Shows the components of an exemplary computing device according to an embodiment of the present invention. Detailed Description
[0012] Due to the ambiguous nature of text-based search queries, users typically spend multiple iterations to fully refine the search query to ensure obtaining all relevant content, but have no money to spend on obtaining irrelevant content (content provider APIs typically deploy a per-volume payment model). This can include envisioning logical items that include relevant keywords and data sources, excluding irrelevant keywords and data sources, and / or considering alternative keywords and / or data sources. Performing this operation in an efficient manner may require expert knowledge. Therefore, there is a need for a simplified method for iteratively searching queries.
[0013] The method for refining an initial search request to a content provider includes the following typical steps: receiving an initial search request from a user and submitting the initial search request to a content provider. The initial search request can be received from the user in a direct manner (e.g., using an input device) or an indirect manner (e.g., forwarded from a computing device via a network to a computing device implementing the method). The content provider may not have to be the same as one of the data sources.
[0014] In response to the submission, the computing device receives a set of sample documents from the content provider. The set of sample documents is accompanied by a source identifier (e.g., URL) that assigns each sample document to the data source that generated the sample document. Within the scope of the present disclosure, a document should be understood as any computer-readable data object carrying human-readable information to be merged by an output device. Without limitation, such human-readable information can be text, image, sound, video, or a combination thereof.
[0015] For simplicity, a data source, which may also be referred to herein as a "source", can be any computing device connected to a network and accessible by using corresponding network calls and routing source identifiers. However, a data source can also be any computing device that can be accessed without a network, for example, by using a local interface of the corresponding computing device, where the local interface of the corresponding computing device, any other component, or a variable stored by the corresponding computing device is identified by the source identifier. Additionally, a data source can be an entity that is not a computing device, including but not limited to an analog information storage, where the source identifier can specify, for example, a non-digital or non-electronic source (such as a book from which a document that can be processed by a computing device implementing the method has been obtained through digitization), a digitizing device (such as a scanner for reproducing information from a non-digital or non-electronic source contained in such a document), and / or a location where the corresponding non-computing source can be found (such as an archive or a library).
[0016] Without limitation, data sources and documents that can be retrieved therefrom can include forums, where the document can be a specific forum, a sub-forum within a larger forum site, or a part of a single page or multiple pages from the corresponding forum or sub-forum; weblogs or blogs, where the document can be a specific blog, a blog section, or a part of a single page or multiple pages from the corresponding blog or blog section; microblogging services, where the document can be a specific microblog belonging to a specific user account or including an excerpt thereof; audio or video download or streaming services, where the document can be a specific audio or video channel or a part of a page including it; review providers, where the document can be a review or discussion, which can be arbitrary or related to a specified product category or product of the provider; virtual social networks, where the document can be a specified public or private user page, a fan page, etc.; or news providers, where the document can be a specified news medium, channel, news on the news medium, or a page representing a part thereof.
[0017] In response to the submission of an initial search request, the amount of sample documents received can be preset by the content provider or specified by the user. In an example, if the total number of search results for an initial search request (e.g., 25 billion) exceeds 1000 documents, the content provider has a rule to limit the set of sample documents to 1000 clicks. Generally, the algorithm incorporating this method can be configured to enhance the set of sample documents by submitting the initial search request to multiple content providers and thus receiving multiple parts of sample documents that are added to the total set of sample documents.
[0018] A computing device applies a topic model to a collection of sample documents to obtain a topic representation. The topic representation is understood herein as a data structure in which keywords derived from the sample documents are organized in a manner that highlights similarities and / or differences between the keywords (e.g., topologically, sequentially, by labels, etc.). Such keywords can be obtained through the statistical analysis included in the topic model, and thus the topic representation can also include weights assigned to the keywords such that each keyword can be used (e.g., processed or displayed) according to its assigned weight. Groups of keywords represented as similar can be understood as the topics described or covered by the sample documents, and keywords represented in different contexts (including the same keyword) can be understood as belonging to different topics. Non-exhaustive examples of topic representations include topic maps, tag clouds or word clouds, tree structures, lists and / or tables of topics, keywords, etc.
[0019] For example, an output device of the computing device presents the topic representation to a user. In response to the presentation, the user inputs a topic relevance score for one or more of the topics included in the topic representation. As used herein, a score can be a continuous measurement, such as a numerical or alphanumeric value, and / or a level or category included in a set of discrete levels or categories.
[0020] The computing device implementing the method classifies data sources based on the topic relevance scores obtained from the user. Thus, one or more of the data sources are assigned a source relevance score that is derived from the (one or more) topic relevance scores input by the user. It must be noted that there is no general limitation on the method that can be used to derive the (one or more) source relevance scores from the (one or more) topic relevance scores, and the examples described herein are given for illustrative purposes only, which may be useful in certain application scenarios, while those skilled in the art can know or envision different other derivation rules that may prove more useful or applicable to other applications.
[0021] The topic relevance score and the source relevance score can be different metrics that must be appropriately mapped. It seems preferable to use the same measurements or categories for the topic and the data source, so such a mapping may be unnecessary. However, if the topic relevance score should correspond to a source relevance score with non-1:1 weighting, then it may be necessary to define a mapping even for the same relevance measurement. In general, a mapping can be defined, for example, that assigns a predefined number of topic relevance levels to a larger number of source relevance levels; maps a given interval of numerical topic relevance scores to a shifted, larger, and / or smaller numerical interval of source relevance scores; converts a category-based topic relevance score to a numerical source relevance score; or converts a numerical topic relevance score to a category-based source relevance score. The topic relevance score and the source relevance score can be designed based on the understanding that relevance can be represented by a one-dimensional measurement that allows for a relative ranking of relevance (e.g., "Document A / Source B is more / less relevant than Document X / Source Y"), and preferably also allows for absolute statements of relevance (e.g., "the most / least relevant document / data source", "the five most / least relevant documents / data sources", "the most / least relevant ten percent of the sample documents / data sources", etc.).
[0022] Based on the source relevance score, the computing device implementing the method determines a refined search request that has increased selectivity for documents covering the highest-scoring topic in the covered topic. The refined search request can be suggested to the user without starting a new search using the refined search request. The algorithm implementing the method can determine more than one refined search request, and these refined search requests can also be presented to the user to select the most appropriate suggestion.
[0023] Due to the one-dimensional nature of the (one or more) relevance scores described above, one or more data sources in the data source can be identified (e.g., using predefined threshold criteria) as having the highest score, which can be understood as a higher score relative to another data source in the data source and / or a higher score relative to a predefined threshold for distinguishing high-relevance sources from low-relevance sources. The increased selectivity can be measured by submitting the refined search request to the content provider, receiving a second set of sample documents in response thereto, counting the number of samples from the second set that cover one or more of the previously identified highest-scoring topics, and comparing that number to the corresponding number of sample documents from the first set that cover one or more of the highest-scoring topics, where the number determined from the second set should be greater than the number determined from the first set.
[0024] Without limitation, selective augmentation can be achieved by adding one or more additional criteria to the initial search request, the one or more additional criteria being reasonably expected to limit the topics covered by the expected second set of sample documents to the previously identified top-scoring topics, and / or to limit the data sources to be included in the expected second search in terms of the previously determined source relevance score(s) to the top-scoring data sources. Without limitation, this can mean including one or more keywords of the top-scoring topics and / or one or more data sources of the top-scoring sources as restrictive or focus criteria in the refined search request, and / or including one or more keywords of the correspondingly defined lowest-scoring topics and / or one or more data sources of the correspondingly defined lowest-scoring sources as exclusion criteria in the refined search request.
[0025] The refined search request can include keywords and / or data sources that will be included in or excluded from the expected search based on the refined search request. Such keywords or data sources can be explicitly selected by the user from the presentation of the topics (which may require including source identifiers in the presentation), or can be automatically derived from the (one or more) topic relevance scores and / or (one or more) source relevance scores assigned by the user, e.g., by including the most frequent (one or more) keywords from each topic marked by the user as relevant or highly relevant; by excluding the most frequent (one or more) keywords from each topic marked by the user as irrelevant or low relevant; by including the (one or more) data sources marked by the user as relevant or highly relevant; and / or by excluding the (one or more) data sources marked by the user as irrelevant or low relevant.
[0026] After determining the refined search request, the user or the algorithm incorporating the method can restart the method by using the refined search request of the completed iteration of the method as the initial search request for the next iteration.
[0027] The method can additionally include, before submitting any search request described herein, querying a content provider for the total number of clicks that would be found if the search request were submitted. In response thereto, the search request can be modified before submission by limiting the count of documents to be retrieved to a number less than the total number of clicks.
[0028] If the number of sample documents retrieved from a data source is less than a predefined minimum number of documents, it is also advantageous to define rules for excluding the data source from the application of the topic model, the presentation of the topic representation, the categorization and / or determination of the refined search request. This can ensure that classifying the data source as highly relevant or lowly relevant is statistically meaningful. For example, if only one sample document is received from a particular data source, its topic relevance may be maximal, such that the data source can be classified with a high source relevance score (the score of the relevant document in the sample is 100%), while other documents from the same data source that would match the initial search request but are not included in the set of sample documents may have a non-maximal topic relevance score, such that if the classification is based on the highly relevant sample document and also the less relevant documents, the classification may have assigned a source relevance score representing lower relevance to the same data source. In this example, by defining a per-source minimum number of documents (e.g., three documents) that must be included in the set of sample documents to consider the data source important enough to obtain a reliable classification, a false high relevance classification can be avoided. Similarly, due to a false low relevance classification, it may not be desirable to exclude the data source.
[0029] Embodiments of the method can have the following advantages: If a second topic representation is determined based on a second set of sample documents retrieved from a content provider in response to submitting a refined search request for another search, during a single iteration of the method, the obtained refined search request is highly selective for documents that cover topics to which the user would assign high relevance. In contrast, manually constructing a refined search request may require multiple iterations of trial and error until the user can identify a suitable combination of keywords and / or data sources to be included in and / or excluded from the search (i.e., a combination that produces search results of the desired relevance). Thus, embodiments of the present invention can reduce the number of search iterations required to obtain search results of the desired relevance. This can reduce the workload of the content provider in searching, accumulating, and delivering clicks (sample documents and / or larger document packages) to the requesting user; can improve the work efficiency of the requesting user; and can reduce the skill level of the user required to construct a complex search request that produces documents of the desired relevance, excluding irrelevant search results but not accidentally excluding relevant search results.
[0030] According to an embodiment, a refined search request includes a logical conjunction of an initial search request and a source criterion, which reduces the number of data sources covered by the refined search request as compared to the number of data sources covered by the initial search request. In other words, the initial search request is extended into an "AND" relationship that connects the initial search request with a statement (i.e., the source criterion) that results in a selection of the data sources reached by the initial search request rather than other data sources, although the initial search request has also reached these data sources, but these data sources will be deselected for future searches based on the refined search request. This can allow focusing such future searches on higher relevance data sources and excluding lower relevance data sources. Thus, a second set of sample documents that can be received during such a future search can include a larger portion of sample documents that have a higher probability of covering high relevance topics compared to the sample documents from the first set of sample documents (which were obtained in response to the initial search request) than the sample documents received from those sources excluded in the refined search request.
[0031] According to one embodiment, the source criterion includes a focus statement for selecting one or more data sources from the data sources based on a source relevance score or an exclusion statement for deselecting one or more data sources from the data sources based on a source relevance score. The focus statement can allow explicitly including documents from data sources with high source relevance scores, while the exclusion statement can allow explicitly excluding documents from data sources with low source relevance scores. Compared to source criteria that implicitly or indirectly reduce the number of covered data sources, this can form a relatively simple way to appropriately reduce the number of covered data sources and can allow targeted specification of data sources to be excluded or included such that indirect source criteria do not accidentally cover low relevance sources.
[0032] According to an embodiment, classification includes selecting a source relevance score from a predefined set of discrete source relevance levels, and the method further includes: for each source relevance level used to determine the source criterion, determining the total number of clicks found for the initial search request at each data source having the corresponding source relevance level; and identifying N data sources having the largest total number of clicks from the data sources having the corresponding source relevance level, where N is a predefined upper limit, and the determination of the refined search request includes, for each source relevance level used to determine the source criterion, populating the source criterion with the N identified data sources.
[0033] The upper limit N can allow for taking into account that many content providers have restrictions on the query length, i.e., they only allow a maximum number of statements that can be included in a single search query. By appropriately choosing N, the refined search query can be designed to ensure that a sufficient number of restrictions can be imposed on the data sources to be covered or excluded, and at the same time, ensure that enough overhead is maintained unaffected by the data source specifications to reserve space for the statements defining the keywords to be included or excluded. For example, if the search provider allows a maximum of 40 statements per query, then N = 30 can be chosen so that up to 30 data sources can be explicitly included or excluded from the refined search request, while still reserving 10 statements for the keywords.
[0034] According to one embodiment, the classification includes selecting a source relevance score from a predefined set of discrete source relevance levels, the determination of the refined search request includes selecting one or more source relevance levels from the source relevance levels for performing a selective increase; and / or the topic relevance score is selected from a predefined set of discrete topic relevance levels.
[0035] This can result in the discretization of the source relevance score and / or the topic relevance score. The discretization of the topic relevance score can simplify the user's decision regarding the assignment of a specific relevance score to a topic, as it may be the case that the user has to assign a numerical score, which may appear more abstract to the user than deciding between multiple discrete options. On the other hand, the discretized source relevance score may be beneficial by avoiding the need for predefined thresholds for determining when a data source has a high or correspondingly low relevance. Such a threshold can ignore the variations in the statistical distribution of the numerical relevance between search requests that produce high scores for relevant sample documents and other search requests that produce low scores for relevant sample documents. In a non-exhaustive example, the set of topic relevance levels includes a high relevance level that allows the user to identify topics that should be kept included in subsequent search queries, and a low relevance level that allows the user to identify topics that should be excluded from subsequent search queries.
[0036] According to an embodiment, the source relevance score and the topic relevance score are selected from a predefined set of discrete common relevance levels, and the determination of the refined search request includes selecting one or more common relevance levels from the common relevance levels for performing a selective increase. This can simplify the mapping of topic relevance levels to source relevance levels, as there is no need to assume a numerical or qualitative correspondence between topic relevance levels and differently defined source relevance levels.
[0037] According to an embodiment, the set of common relevance levels includes a high relevance level, and the classification includes: if the relative frequency of the sample documents that are associated with a given data source and cover the topics assigned to the high relevance level is equal to or exceeds a predefined high relevance threshold, then the given data source is assigned to the high relevance level. The determination of the refined search request includes: in the case where the high relevance level is used for the determination of the refined search request, compared with the initial search request, the refined search request is restricted to one or more data sources in the data sources that are assigned to the high relevance level. According to one embodiment, the high relevance threshold is 100%.
[0038] The high relevance level can be considered to represent documents and data sources that have the highest relevance compared to any other relevance level in the set of common relevance levels. Specifically, but not necessarily, the high relevance threshold can be set to 100%, in which case, the data sources that only provide documents covering (high) relevant topics can be classified as the high relevance level. Otherwise, such data sources can also be included under the high relevance level that provides no more than a predefined portion of the documents covering less relevant topics. This relaxation of the high relevance filtering can increase the total number of documents considered to be high relevance documents, which can result in more diverse topics and / or document types obtained for subsequent analysis.
[0039] According to an embodiment, the set of common relevance levels further includes a low relevance level, and the classification includes: if the relative frequency of the sample documents that are associated with a given data source and cover the topics assigned to the low relevance level is equal to or exceeds a predefined first low relevance threshold, then the given data source is assigned to the low relevance level, and / or if the relative frequency of the sample documents that are associated with a given data source and cover the topics assigned to the high relevance level is equal to or less than a predefined second low relevance threshold, then the given data source is assigned to the low relevance level. The determination of the refined search request includes: in the case where the low relevance level is used for the determination of the refined search request, compared with the initial search request, one or more data sources in the data sources that are assigned to the low relevance level are excluded from the refined search request. According to an embodiment, the first low relevance threshold is 100% and / or the second low relevance threshold is 0%.
[0040] A low relevance level can be considered to represent, respectively, a document and a data source that have the lowest relevance compared to any other relevance level in a set of common relevance levels. Specifically, but not necessarily, a first low relevance threshold can be set to 100%, in which case, a data source that provides only documents covering low-relevant topics can be classified as a low relevance level. Additionally or alternatively, and still not necessarily, a second low relevance threshold can be set to 0%, in which case, a data source that does not provide documents covering high-relevant topics can be classified as a low relevance level. Otherwise, when these conditions are respectively relaxed to values below 100% for the first low relevance threshold, or above 0% for the second low relevance threshold, such a data source can also be included under a low relevance level that provides no more than a predefined portion of documents covering topics of higher relevance, and / or such a data source can be included under a low relevance level that provides no more than a predefined portion of documents covering topics of higher relevance. This relaxation of the low relevance filtering can increase the total number of documents considered to be low relevance documents, which can result in a more comprehensive exclusion of data sources with a low return of high relevant documents.
[0041] The set of common relevance levels can include additional relevance levels, such as a medium relevance or a partial relevance level, which can be assigned to topics and / or data sources that cannot be assigned to either a high relevance level or a low relevance level.
[0042] According to an embodiment, the method further includes obtaining a precision level, the refined search request is further determined based on the precision level, the precision level is obtained from a predefined set of discrete precision levels, and the selective increase is based on the precision level. The precision level can reflect the need to create a search query that balances precision (ensuring that no computing resources and money are spent on obtaining irrelevant content) and recall (ensuring that all relevant content is obtained). Incorporating the precision level can allow the decision on which source relevance level to use to determine the refined search request to be linked to user preferences and / or the statistical properties of a set of sample documents, and thus can simplify the determination of the refined search request. Without limitation, the precision level can be received as user input (e.g., using a software-implemented slider or radio button field), or can be determined by the computing device executing the method by analyzing the set of received sample documents.
[0043] According to an embodiment, the precision level is obtained as an input from the user. This can give the user additional freedom to influence the result of the method according to the precision requirements of the individual situation, which can be dedicated to a more advanced task of the user involving the refinement of the search request and can violate any statistical findings, such as the keyword frequencies that the computing device implementing the method can find in the set of received sample documents.
[0044] According to an embodiment, the set of precision levels includes a high precision level, a balanced precision level, and a high recall level, the source relevance score is selected from the group consisting of: a high relevance level, a medium relevance level, and a low relevance level, and the classification includes assigning a given data source to the high relevance level if the relative frequency of sample documents associated with the given data source and covering the topics assigned to the high relevance level is equal to or exceeds a predefined high relevance threshold, the classification further includes assigning the given data source to the low relevance level if the relative frequency of sample documents associated with the given data source and covering the topics assigned to the low relevance level is equal to or exceeds a predefined first low relevance threshold, and / or if the relative frequency of sample documents associated with the given data source and covering the topics assigned to the high relevance level is equal to or less than a predefined second low relevance threshold, and the classification further includes assigning the data source to the medium relevance level if the given data source cannot be assigned to either the high relevance level or the low relevance level. In the case where the obtained precision level is the high precision level, the refined search request includes a logical conjunction of the initial search request and a focus statement that selects one or more data sources having the high relevance level as the source relevance score from the data sources. In the case where the obtained precision level is the balanced precision level, the refined search request includes a logical conjunction of the initial search request and a focus statement that includes a logical disjunction of one or more data sources having the high relevance level as the source relevance score from the data sources and a logical disjunction of one or more data sources having the medium relevance level as the source relevance score from the data sources. In the case where the obtained precision level is the high recall level, the refined search request includes a logical conjunction of the initial search request and a focus statement that includes a logical negation of one or more data sources having the low relevance level as the source relevance score from the data sources.
[0045] This can constitute a favorable correspondence between the set of precision levels and the source relevance score, which can simplify the selection of which source relevance level to use to determine the refined search request. Let K be the initial search request, S R be the set of relevant data sources (i.e., the set of data sources classified with the high relevance level), S P be the set of partially relevant data sources (i.e., the set of data sources classified with the medium relevance level), and S I be the set of irrelevant data sources (i.e., the set of data sources classified with the low relevance level). Then, the three cases for determining the refined search request described above can be briefly written for the high precision level as "K AND S R”, i.e., determining a refined search request by restricting the initial search request to highly relevant data sources; for a balanced precision level, it is abbreviated as “K AND (S R OR S P )”, i.e., determining the refined search request by restricting the initial search request to also highly relevant data sources and partially relevant data sources; and for high recall level, it is abbreviated as “K AND (NOT S I )”, that is, determining a refined search request by excluding irrelevant data sources from the initial search request.
[0046] It can be seen that the refined search request determined for the high precision level can produce the strictest focus on highly relevant documents, while the balanced precision level can include additional highly relevant documents at the expense of including an additional portion of documents that may be outside the high relevance level (i.e., may cover topics of lover relevance). It can be further noted that both the high precision level and the balanced precision level correspond to a strategy of focus refinement, i.e., the spectrum of data sources is explicitly limited to sources with a high probability of delivering highly relevant documents, while the high recall level corresponds to a strategy of exclusion refinement, i.e., only the least relevant data sources are explicitly excluded, so that a more diverse range of high and medium relevance topics can be expected for the high recall case. Therefore, if the number of highly relevant search results can be expected to be small compared to the number of irrelevant documents, the high recall level can be recommended, so that excluding low-relevance data sources can maximize the number of relevant search results. On the other hand, if the number of relevant search results is large compared to the number of irrelevant documents, and / or if the initial search request is highly ambiguous, the high precision level can be recommended, so that high selectivity can be expected to filter out all topics or keywords with less relevance. Furthermore, it may be advantageous to have a balanced accuracy level as a standard setting as a suggested option to the user, which may be helpful if neither a high precision level nor a high recall level is clearly advantageous.
[0047] According to an embodiment, the topic relevance score is selected from a topic relevance group including a high relevance level, the high relevance level representing the highest relevance compared to all other relevance levels in the topic relevance group, and the method further comprises determining a number D of sample documents covering the topic having the high relevance level as the topic relevance score. R If D R The ratio D of the total number of sample documents in the set of sample documents R / D is less than or equal to the predefined lower precision threshold, the high precision level is set to the precision level. RIf D is greater than the lower precision threshold and less than the predefined upper precision threshold, then set the balanced precision level to the precision level. If D R is greater than or equal to the upper precision threshold, then set the high recall level to the precision level.
[0048] In this way, the precision type can be automatically determined by statistical analysis of the set of received sample documents. It can be noted that only the number D R of highly relevant sample documents and the total number D of sample documents are required to obtain a decision for determining the precision level of the refined search request. Additionally, it may not be necessary to retrieve further information from the content provider to make this decision. In this way, the refined search request can be determined with minimal interaction with the user and the content provider and with efficient use of the computing resources of the computing device implementing the method.
[0049] According to an embodiment, the topic relevance score is selected from a group of topic relevance levels including a high relevance level, where the high relevance level represents the highest relevance compared to all other relevance levels in the group of topic relevance levels. The method further includes querying the content provider for the number D RS of available documents as search results for a hypothetical search request, where the hypothetical search request includes a logical combination of an initial search request and a focus statement that selects one or more data sources in the data source with a high relevance level as the source relevance score. The method further includes querying the content provider for the number D K of available documents as search results for the initial search request. RS If the ratio of D K to D RS D K / D is less than or equal to the lower precision threshold, then set the high precision level to the precision level. If D RS / D K is greater than the lower precision threshold and less than the upper precision threshold, then set the balanced precision level to the precision level. If D RS / D K is greater than or equal to the upper precision threshold, then set the high recall level to the precision level.
[0050] If the total number of clicks D KIf it is larger than the number D of received sample documents, then this way of determining the precision level for managing the refined search request can be particularly advantageous, but not necessary. In this case, the set of sample documents may be too small to carry statistical relevance, so a more accurate decision can be made if the statistics of the entire click are considered instead of only considering the statistics of the set of sample documents. However, if it is desired to deliver a high-quality refined search request by minimizing the number of assumptions used to execute the method, then making a precision level decision based on not considering the ratio of D K to D RS and D K is also advantageous.
[0051] According to an embodiment, the method further includes: in response to the presentation, receiving, from the user, a document relevance score for one or more of the sample documents in the sample documents, and further classifying the data source based on the document relevance score; and / or receiving, from the user, a source relevance score for one or more of the data sources in the data sources, and exempting the data source with the source relevance score received from the user from classification; and / or receiving, from the user, a keyword relevance score for one or more keywords representing one or more of the topics in the keywords, and further classifying the data source based on the keyword relevance score.
[0052] The possibility of providing the user with a relevance score for an item other than the topic can increase the probability of obtaining a refined search request with high selectivity for highly relevant documents, even if the user is unsure which relevance to assign to some topics identified by the topic model, or if the identified topics do not seem to provide appropriate support for such relevance expected by the user. For example, the user may find a specific sample document that is more relevant than all the presented topics and can accordingly mark the document as highly relevant. Then, the algorithm implementing the method can, for example, obtain the desired increase in selectivity by focusing the refined search request on the data source from which the sample document marked as highly relevant was obtained, and / or focusing the refined search request on the feature keywords of the sample document. Similarly, the refined search request can be configured to exclude keywords and / or sources related to irrelevant documents; include data sources specified by the user as highly relevant and / or exclude data sources specified by the user as irrelevant; and / or include keywords specified by the user as highly relevant and / or exclude keywords specified by the user as irrelevant.
[0053] According to an embodiment, the method further includes: applying a predefined default topic relevance score to a topic for which no topic relevance score is received; and / or applying a predefined default source relevance score to a data source for which its source relevance score is not determined during classification. This can maximize the number of accessible topics and / or data sources as a basis for determining the refined search request.
[0054] According to an embodiment, topics are presented that are restricted to be equal to or exceed a predefined minimum number of clicks for an initial search request. This can prevent users from marking as relevant or irrelevant topics with low statistical significance due to too few clicks on the keywords within those topics. Similarly, the presentation of statistically non-significant data sources, keywords, and / or sample documents can be suppressed to prevent users from accidentally marking such data sources, keywords, and / or sample documents as relevant or irrelevant.
[0055] Turning now to the drawings, Figure 1 An exemplary routine is shown for processing an initial search request K in response to an initial search request K using an exemplary computing environment adapted to retrieve sample documents. User 101 uses a computing device 10 that is connected via a computing network 120 to another computing device (e.g., a server) of a document provider 130. The user inputs the initial search request K into the computing device 10 and submits the initial search request to the content provider 130 via the network 120. Without limitation, the initial search request K can include keywords to be included or excluded in the requested search and / or source identifiers (e.g., URLs) of data sources to be included or excluded in the requested search. The initial search request K can also include an identifier that indicates to the content provider 130 that the requested search should be restricted to a set of sample documents, where the specification of the set of sample documents to be returned (such as the maximum number of clicks to be included in the set of sample documents) can be predetermined by the content provider 130 and / or by a corresponding specification also included in the initial search request K.
[0056] In response, the content provider 130 parses the initial search request K and, if the initial search request K is valid, performs a search for documents that match the initial search request K. According to the given specification of the set of sample documents to be returned, the document provider 130 aggregates the set of sample documents 132 and delivers the set 132 to the computing device 10 of the user 101 via the network 120.
[0057] Figure 2 Shows Figure 2 An exemplary topic representation 200 formed by a list of topics 202, 204, 206, 208, and 210 in the non-limiting specific case shown in, where each topic is represented by a list of keywords printed in the style of a tag cloud, i.e., having different font sizes corresponding to the relative frequency of the respective keywords within the documents covering the respective topics. Figure 2The examples shown in FIG. may be used to assign topic relevance scores to one or more of the topics 202, 204, 206, 208, and 210, and also optionally to assign keyword relevance scores to one or more of the keywords shown. The particular choice of the list of label-cloud topics as the topic representation 200 is merely illustrative; without limitation, other choices of the topic representation 200 may utilize topic maps, label clouds, tree structures, lists, tables, or combinations thereof.
[0058] Figure 3 FIG. is a diagram showing the relationships that can facilitate the processing of relevance scores according to an embodiment of the present invention. By applying a topic model that is assumed to be known and thus not described herein, a set 132 of received sample documents is analyzed and individual sample documents are assigned to topics. The topics are presented to the user 101 as the topic representation 200 including the interface 300, thereby allowing the user 101 to set the topic relevance scores for one or more of the topics. In Figure 3 In the non-exhaustive examples shown in FIG., each topic may be assigned one topic relevance level from a group including the topic relevance levels "relevant", "partially relevant", and "irrelevant". In Figure 3 In a specific instance of FIG., the topic relevance level "relevant", which may also be referred to as a high relevance level, is assigned to topic 1; the topic relevance level "partially relevant", which may also be referred to as a medium relevance level, is assigned to topic 2; and the topic relevance level "irrelevant", which may also be referred to as a low relevance level, is assigned to topic p.
[0059] Figure 3The illustration is simplified by showing only four data sources 310, from which six sample documents in the collection 132 are derived. Topic 1 is populated by one document from source S1 and one document from source Sn; Topic 2 is populated by one document from source S1; and Topic p is populated by one document from source S2, one document from source S3, and one document from source Sn. Thus, source S1 produces one highly relevant sample document and one partially relevant document; sources S2 and S3 each produce one low-relevant sample document; and source Sn produces one highly relevant sample document and one low-relevant document. In the simplified example shown, an exemplary and purely illustrative way to classify the data sources 310 as source relevance scores could be to assign weights to the topic relevance scores, e.g., 100% for each document with high relevance, 50% for each document with partial relevance, and 0% for each document with low relevance. In the same example, the average for each data source could produce a weight of 75% for source S1, 0% for sources S2 and S3, and 50% for source Sn. Still in a non-limiting example, a threshold of 33% average weight could be applied to decide between a low source relevance level and a medium source relevance level, and a threshold of 67% average weight could be applied to decide between a high source relevance level and a medium source relevance level. In the same example, this could result in classifying S1 as a high-relevance data source; classifying S2 and S3 as low-relevance data sources; and classifying Sn as a medium-relevance data source. It must be noted that the present invention anticipates no limitation to the application of any modified or different method for deriving source relevance scores from topic relevance scores.
[0060] Figure 4 is a flowchart showing the steps of an exemplary method for refining an initial search request to a content provider according to an embodiment of the present invention. The steps of the method are described below from the perspective of a computing device 10 implementing the algorithm of the method. The method includes step 402, where an initial search request is received from a user 101. Subsequently, at step 404, the initial search request is submitted to a content provider 130, e.g., using a computing or communication network 120. Then, at step 406, a collection 132 of sample documents matching the initial search request is received from the content provider 130. A data source identifier representing a data source 310 is assigned to each sample document, from which the corresponding sample document is derived.
[0061] At step 408, a topic model is applied to the set 132 of received sample documents. The topic model can identify mutually exclusive topics and can assign each of the sample documents in the sample document set to one of the identified topics. Step 408 also includes obtaining a topic representation 200 of the topics that can be output to a user. Then, at step 410, the topic representation is presented to the user, for example, using an output device of the computing device 10. In response to the presentation at step 410, at step 412, a topic relevance score for one or more of the presented topics is received from the user.
[0062] Next, at step 414, the data source 310 is classified based on the topic relevance scores. Accordingly, one or more of the data sources in the data source 310 are assigned a source relevance score that is derived from the received (one or more) topic relevance scores by applying appropriate logic. Subsequently, at step 416, further logic is applied based on the obtained (one or more) source relevance scores to determine a refined search request. The refined search request may be based on the initial search request or may be newly constructed based on keywords, topic relevance scores, and / or source relevance scores obtained during steps 408, 412, and 414. Compared to the initial search request, the refined search request is constructed in such a way that it has increased selectivity for sample documents in topics that have been assigned the highest topic relevance scores during the user assignment at step 412, and / or for other documents that match the initial search request and would be assigned to one or more of these highest-scoring topics if such other documents were included in the sample document set 132.
[0063] Embodiments of the present invention may be implemented using a computing device, which may also be referred to as a computer system, a client, or a server. Now refer to Figure 5 , a schematic diagram showing an example of a computer system is illustrated. The computer system 10 is only one example of a suitable computer system and is not intended to impose any limitation on the scope of use or functionality of the embodiments of the present invention described herein. In any case, the computer system 10 is capable of implementing and / or performing any of the functions set forth above.
[0064] In computer system 10, there is a computer system / server 12 that operates with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations that can be applicable for use with computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or notebook devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed computing environments including any of the above systems or devices, etc.
[0065] Computer system / server 12 can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally speaking, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform specific tasks or implement specific abstract data types. Computer system / server 12 can be implemented in a distributed computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed computing environment, program modules can be located in local and remote computer system storage media including memory storage devices.
[0066] As Figure 5 shown, the computer system / server 12 in computer system 10 is shown in the form of a general-purpose computing device. The components of computer system / server 12 may include, but are not limited to, one or more processors or processing units 16, a system memory 28, and a bus 18 that couples various system components including the system memory 28 to the processor 16. Bus 18 represents any one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any one of various bus architectures. By way of example and not limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0067] Computer system / server 12 typically includes various computer system-readable media. Such media can be any available media accessible by computer system / server 12, and it includes volatile and non-volatile media, removable and non-removable media.
[0068] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer system / server 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, there is provided a storage system for reading from and writing to a non-removable non-volatile magnetic medium (not shown and typically referred to as a "hard disk drive"). Although not shown, a disk drive for reading from and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable non-volatile optical disk such as a CD-ROM, DVD-ROM or other optical media may be provided. In such instances, each external device may be connected to the bus 18 via one or more data media interfaces. As will be further shown and described below, memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to carry out the functions of embodiments of the present invention.
[0069] A program / utility 40 having a set (at least one) of program modules 42, as well as an operating system, one or more application programs, other program modules, and program data, may be stored in memory 28 by way of example and not limitation. Each or some combination of the operating system, one or more application programs, other program modules, and program data may include an implementation of a network environment. The program modules 42 generally carry out the functions and / or methods of embodiments of the present invention as described herein.
[0070] The computer system / server 12 may also communicate with one or more external devices 14 such as a keyboard, a pointing device, a display 24, etc.; communicate with one or more devices that enable a user to interact with the computer system / server 12; and / or communicate with any device that enables the computer system / server 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may occur via an input / output (I / O) interface 22. Additionally, the computer system / server 12 may communicate via a network adapter 20 with one or more networks such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet). As shown, the network adapter 20 communicates with other components of the computer system / server 12 via the bus 18. It should be understood that although not shown, other hardware and / or software components may be used in conjunction with the computer system / server 12. Examples include, but are not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.
[0071] A computer system, such as Figure 5The computer system 10) shown in can be used to perform the operations disclosed herein, such as refining an initial search request for a content provider. Such a computer system can be a stand-alone computer without a network connection, which can receive data to be processed, such as an initial search request, a collection of sample documents with corresponding source identifiers, or a topic relevance score, through a local interface. However, a computer system connected to a network (such as, a communication network and / or a computing network) can also be used to perform such operations.
[0072] The present invention can be a system, method, and / or computer program product at any possible integrated technical detail level. The computer program product can include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention.
[0073] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punch card or a raised structure in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable) or an electrical signal transmitted through a wire.
[0074] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device via a network (such as, the Internet, a local area network, a wide area network, and / or a wireless network), or to an external computer or an external storage device. The network can include a copper transmission cable, an optical transmission fiber, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0075] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++) and procedural programming languages (such as the C programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), can execute the computer-readable program instructions by personalizing the electronic circuit using the state information of the computer-readable program instructions to perform aspects of the present invention.
[0076] Aspects of the present invention will be described herein with reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0077] These computer-readable program instructions can be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create a means for implementing the functions / actions specified in the flowchart and / or block diagram block or blocks. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, so that the computer-readable storage medium storing the instructions includes an article of manufacture containing instructions that implement aspects of the functions / actions specified in the flowchart and / or block diagram block or blocks.
[0078] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in the flowchart and / or block diagram block or blocks.
[0079] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of the possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified (one or more) logical functions. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be completed as one step, executed simultaneously, substantially simultaneously, partially or fully in time-overlapped manners, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or combinations of special purpose hardware and computer instructions.
[0080] It should be understood that although this disclosure includes a detailed description of cloud computing, the implementations of the teachings recited herein are not limited to cloud computing environments. On the contrary, embodiments of the present invention are capable of being implemented in conjunction with any other type of computing environment now known or later developed.
Claims
1. A computer-implemented method for refining an initial search request for a content provider, the method comprising: Receiving an initial search request from a user; Submitting the initial search request to a content provider; Receiving from the content provider a collection of sample documents and source identifiers of corresponding sample documents in the sample documents, the source identifiers identifying corresponding data sources in a data source that are associated with the corresponding sample documents in the sample documents; Applying a topic model to the collection of sample documents to obtain a topic representation, the topic representation being a description of the topics covered by the corresponding sample documents in the sample documents; Presenting the topic representation to the user; Receiving from the user a topic relevance score for a corresponding topic in the topics; Classifying the data sources according to the topic relevance scores to obtain source relevance scores for the corresponding data sources in the data source; And Based on the source relevance scores, determining a refined search request that has increased selectivity for documents covering the highest-scoring topic in the topics.
2. The computer-implemented method according to claim 1, wherein the refined search request comprises a logical combination of the initial search request and a source criterion, and wherein the source criterion reduces the number of data sources covered by the refined search request compared to the number of data sources covered by the initial search request.
3. The computer-implemented method according to claim 2, wherein the source criteria comprise: Selecting focus statements of one or more data sources in the data source based on the source relevance scores, or deselecting exclusion statements of one or more data sources in the data source based on the source relevance scores.
4. The computer-implemented method according to any one of claims 2 to 3, wherein classifying the data sources comprises selecting a source relevance score from a predefined set of discrete source relevance levels, and wherein the computer-implemented method further comprises: For a corresponding discrete source relevance level in the discrete source relevance levels used to determine the source criterion, determining the total number of clicks found for the initial search request at the corresponding data source in the data source having the corresponding discrete source relevance level in the discrete source relevance levels; And Identifying N data sources in the data source having the corresponding discrete source relevance level in the discrete source relevance levels that have the largest total number of clicks, where N is a predefined upper limit, and wherein determining the refined search request comprises: for a corresponding discrete source relevance level in the discrete source relevance levels used to determine the source criterion, populating the source criterion with the N identified data sources.
5. The computer-implemented method according to any one of claims 1 to 3, wherein classifying the data sources comprises selecting a source relevance score from a predefined set of discrete source relevance levels, and wherein determining the refined search request comprises selecting one or more source relevance levels in the source relevance levels for performing an increase in selectivity, and wherein the topic relevance scores are selected from a predefined set of discrete topic relevance levels.
6. The computer-implemented method according to claim 1, wherein the source correlation score and the topic correlation score are selected from a predefined set of discrete common correlation levels, and wherein determining the refined search request comprises: Selecting one or more discrete common relevance levels in the discrete common relevance levels for performing an increase in selectivity.
7. The computer-implemented method according to claim 6, wherein the predefined set of discrete common relevance levels includes a high relevance level, and wherein classifying the data source comprises: If the relative frequency of sample documents that are associated with a given data source and cover topics assigned to a high relevance level is equal to or exceeds a predefined high relevance threshold, then the given data source is assigned to the high relevance level, where determining a refined search request includes: in the case where the high relevance level is used to determine the refined search request, restricting the refined search request to one or more data sources assigned to the high relevance level in the data source, as compared to the initial search request.
8. The computer-implemented method according to claim 7, wherein the predefined high relevance threshold is one hundred percent.
9. The computer-implemented method according to any one of claims 6 to 8, wherein the predefined set of discrete common relevance levels further includes a low relevance level, and wherein classifying the data source comprises: If the relative frequency of sample documents that are associated with a given data source and cover topics assigned to a low relevance level is equal to or exceeds a predefined first low relevance threshold, and / or if the relative frequency of sample documents that are associated with a given data source and cover topics assigned to a high relevance level is equal to or less than a predefined second low relevance threshold, then the given data source is assigned to the low relevance level, where determining a refined search request includes: in the case where the low relevance level is used to determine the refined search request, excluding from the refined search request one or more data sources assigned to the low relevance level in the data source, as compared to the initial search request.
10. The computer-implemented method according to claim 9, wherein the first low relevance threshold is one hundred percent and the second low relevance threshold is zero percent.
11. The computer-implemented method according to claim 1, the method further comprising: obtaining a precision level, where the refined search request is further determined based on the precision level, where the precision level is obtained from a predefined set of discrete precision levels, and where an optional increase is based on the precision level.
12. The computer-implemented method according to claim 11, wherein the precision level is obtained as an input from a user.
13. The computer-implemented method according to any one of claims 11 to 12, wherein the set of precision levels includes a high precision level, a balanced precision level, and a high recall level. wherein the corresponding source correlation score in the source correlation score is selected from the group consisting of a high correlation level, a medium correlation level, and a low correlation level, and wherein classifying the data source includes: If the relative frequency of the sample documents that are associated with a given data source and cover topics assigned to a high relevance level is equal to or exceeds a predefined high relevance threshold, then the given data source is assigned to a high relevance level, wherein classifying the data source further includes: if the relative frequency of the sample documents that are associated with the given data source and cover topics assigned to a low relevance level is equal to or exceeds a predefined first low relevance threshold, and / or if the relative frequency of the sample documents that are associated with the given data source and cover topics assigned to a high relevance level is equal to or less than a predefined second low relevance threshold, then the given data source is assigned to a low relevance level, wherein classifying the data source further includes: if the given data source cannot be assigned to either a high relevance level or a low relevance level, then the given data source is assigned to a medium relevance level, wherein in a case where the precision level is a high precision level, the refined search request includes a logical conjunction of an initial search request and a focus statement that selects one or more data sources having a high relevance level as a source relevance score from the data sources, wherein in a case where the precision level is a balanced precision level, the refined search request includes a logical conjunction of an initial search request and a focus statement that includes a logical disjunction of one or more data sources having a high relevance level as a source relevance score from the data sources and a logical disjunction of one or more data sources having a medium relevance level as a source relevance score from the data sources, wherein, in a case where the precision level is a high recall level, the refined search request includes a logical conjunction of an initial search request and a focus statement that includes a logical negation of one or more data sources having a low relevance level as a source relevance score from the data sources.
14. The computer-implemented method according to claim 13, wherein the respective topic relevance scores in the topic relevance scores are selected from a topic relevance group that includes a high relevance level representing the highest relevance compared to all other relevance levels in the topic relevance group, wherein the computer-implemented method further includes: Determine the number D of sample documents that cover the topics with a high relevance level as the topic relevance score R , where if D R The ratio D of D to the total number D of sample documents in the set of sample documents R / D is less than or equal to a predefined lower precision threshold, then set the high precision level as the precision level; where if D R / D is greater than the predefined lower precision threshold and less than the predefined upper precision threshold, then set the balanced precision level as the precision level, and where if D R / D is greater than or equal to the predefined upper precision threshold, then set the high recall level as the precision level.
15. The computer-implemented method according to claim 13, wherein the respective topic relevance scores in the topic relevance scores are selected from a topic relevance group that includes a high relevance level representing the highest relevance compared to all other relevance levels in the topic relevance group, wherein the method further includes: Query a content provider for the number D of available documents that are search results for a hypothetical search request RS , where the hypothetical search request includes a logical combination of an initial search request and a focus statement that selects one or more data sources in the data source that have a high relevance level as a source relevance score; and Query the content provider for the number D of available documents that are search results of the initial search request K , where if D RS is in ratio to D K such that D RS / D K is less than or equal to a predefined lower precision threshold, then set the high precision level as the precision level, where if D RS / D K is greater than the predefined lower precision threshold and less than a predefined upper precision threshold, then set the balanced precision level as the precision level, where if D RS / D K is greater than or equal to the predefined upper precision threshold, then set the high recall level as the precision level.
16. The computer-implemented method according to any one of claims 1 to 3, the method further includes: receiving from the user a document relevance score of the respective sample document in the sample documents, and classifying the data source is further based on the document relevance score; receiving from the user a source relevance score of the respective data source in the data sources, and a data source having a source relevance score received from the user is exempt from classifying the data source; and Receiving, from a user, a keyword relevance score for a corresponding keyword in a keyword that represents a corresponding topic in a topic, and classifying the data source further based on the keyword relevance score.
17. The computer-implemented method according to any one of claims 1 to 3, the method further comprising: Applying a predefined default topic relevance score to a topic for which no topic relevance score has been received; And Applying a predefined default source relevance score to a data source for which a source relevance score has not been determined during classification of the data source.
18. The computer-implemented method according to any one of claims 1 to 3, wherein the topic representation is restricted to topics that are equal to or exceed a predefined minimum number of clicks for an initial search request.
19. A computer program product for refining an initial search request to a content provider, the computer program product comprising a computer-readable storage medium having program instructions implemented therein, the program instructions executable by one or more processors, the program instructions executable to: Receive an initial search request from a user; Submit the initial search request to a content provider; Receive, from the content provider, a set of sample documents and source identifiers for corresponding sample documents in the sample documents, the source identifiers identifying corresponding data sources in the data source associated with the corresponding sample documents in the sample documents; Apply a topic model to the set of sample documents to obtain a topic representation, the topic representation being a description of the topics covered by the corresponding sample documents in the sample documents; Present the topic representation to the user; Receive, from the user, a topic relevance score for a corresponding topic in the topics; Classify the data source according to the topic relevance score to obtain a source relevance score for the corresponding data source in the data source; And Based on the source relevance score, determine a refined search request that has increased selectivity for documents covering the topic with the highest score in the topics.
20. A computer system for refining an initial search request to a content provider, the computer system comprising: One or more processors, one or more computer-readable tangible storage devices, and program instructions stored in at least one of the one or more computer-readable tangible storage devices for execution by at least one of the one or more processors, the program instructions executable to: Receive an initial search request from a user; Submit the initial search request to a content provider; Receive, from the content provider, a set of sample documents and source identifiers for corresponding sample documents in the sample documents, the source identifiers identifying corresponding data sources in the data source associated with the corresponding sample documents in the sample documents; Apply a topic model to the set of sample documents to obtain a topic representation, the topic representation being a description of the topics covered by the corresponding sample documents in the sample documents; Present the topic representation to the user; Receive, from the user, a topic relevance score for a corresponding topic in the topics; Classify the data source according to the topic relevance score to obtain a source relevance score for the corresponding data source in the data source; And Based on the source relevance score, a refined search request is determined, and the refined search request has increased selectivity for documents covering the topic with the highest score among the topics.
Citation Information
Patent Citations
Query representation and hybrid retrieval model construction method based on context sensing theme
CN106294662A
Identifying inadequate search content
US9020933B2