Improving search quality based on keyword search and contextual search
The integration of keyword and contextual vector-based search results with AHP weighting improves search accuracy and relevance, addressing inaccuracies in existing methods, especially in small datasets.
Patent Information
- Application Number
- PCT/CN2024/100893
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-24
- Publication Date
- 2026-01-02
AI Technical Summary
Existing keyword and contextual search methods suffer from inaccuracies due to misinterpretation of context and reliance on exact matches, leading to irrelevant results, especially in small datasets.
A method that combines keyword and contextual vector-based search results, using weights determined by analytic hierarchy process (AHP) to rank and combine search results, improving accuracy and relevance.
Enhances search result quality by integrating keyword and contextual relevance, providing more accurate and diverse search outcomes, particularly in small datasets.
Smart Images

Figure CN2024100893_02012026_PF_FP_ABST
Abstract
Description
IMPROVING SEARCH QUALITY BASED ON KEYWORD SEARCH AND CONTEXTUAL SEARCHTechnical Field
[0001] Various examples of the disclosure generally relate to executing a search query, e.g., by a database management system. Various examples specifically relate to determining a set of search results based on a first set of search results obtained based on one or more keywords associated with a search query and a second set of search results obtained based on one or more contextual vectors associated with the search query.Background
[0002] Information retrieval is the task of identifying and retrieving information system resources that are relevant to an information need. The information need can be specified in the form of a search query. In the case of document retrieval, queries can be based on full-text or other content-based indexing. Information retrieval comprises searching for information in a document, searching for documents themselves, and also searching for the metadata that describes data, and for databases of texts, images, videos, or sounds. An information retrieval system is a software system that provides access to books, journals, videos, and other documents; it also stores and manages those documents.
[0003] An information retrieval process begins when a user enters a search query into an information retrieval system. Queries are statements of information needs, for example, search strings in web search engines. A query does not uniquely identify a single object in the collection. Instead, several objects may match the query, perhaps with different degrees of relevance.
[0004] An object is an entity that is represented by information in a content collection or database. Depending on the application the data objects may be, for example, text documents, images, audio, or videos. Often the documents themselves are not kept or stored directly in the information retrieval system but are instead represented in the system by document surrogates or metadata.
[0005] Keyword search is a widely used method for information retrieval. For example, some web search engines are designed to search for keywords in a document, e.g., the title, the body, and so on. A keyword can be any term that exists within the document. Keywords can be further refined using Boolean operators such as "AND, OR, NOT. " However, keyword search has several drawbacks. For example, one or more keywords can have different meanings, causing confusion and leading to irrelevant search results, which reduces the accuracy of search results. Keyword searches often ignore the context in which the keyword is used, leading to results that are not relevant to the user's intent. Keyword searches require an exact match, which means that slight variations or misspellings in the search term can lead to missed results.
[0006] Contextual search aims to improve upon keyword search by considering the context of the query to deliver more relevant results. However, it also has its drawbacks. For example, contextual search systems can misinterpret the user's intent or the context, for example, if the search query comprises one or two words, leading to irrelevant or incorrect results. The quality of the contextual search is heavily dependent on the quality and quantity of the data it is trained on, which makes it challenging to apply contextual search to a small dataset or database.Summary
[0007] Therefore, a need exists for advanced techniques for improving search quality. Specifically, a need exists for advanced techniques of automatically determining based on a search query a set of search results with improved accuracy and liability.
[0008] This need is met by the features of the independent claims. The features of the dependent claims define embodiments.
[0009] A computer-implemented method is provided. The method comprises obtaining a first set of search results based on one or more keywords associated with a search query and obtaining a second set of search results based on one or more contextual vectors associated with the search query. The method also comprises determining a first weight associated with the first set of search results and a second weight associated with the second set of search results. The first weight and the second weight respectively represent how well the first set of search results and the second set of search results match the search query. The method further comprises determining a set of search results based on the first weight and the second weight.
[0010] A computer program product or a computer program or a computer-readable storage medium including program code is provided. The program code can be executed by at least one processor. Executing the program code causes the at least one processor to perform a method. The method comprises obtaining a first set of search results based on one or more keywords associated with a search query and obtaining a second set of search results based on one or more contextual vectors associated with the search query. The method also comprises determining a first weight associated with the first set of search results and a second weight associated with the second set of search results. The first weight and the second weight respectively represent how well the first set of search results and the second set of search results match the search query. The method further comprises determining a set of search results based on the first weight and the second weight.
[0011] A computing or processing device comprising at least one processor and a memory is provided. Upon loading and executing program code from the memory, the at least one processor is configured to perform a method. The method comprises obtaining a first set of search results based on one or more keywords associated with a search query and obtaining a second set of search results based on one or more contextual vectors associated with the search query. The method also comprises determining a first weight associated with the first set of search results and a second weight associated with the second set of search results. The first weight and the second weight respectively represent how well the first set of search results and the second set of search results match the search query. The method further comprises determining a set of search results based on the first weight and the second weight.
[0012] It is to be understood that the features mentioned above and those yet to be explained below may be used not only in the respective combinations indicated, but also in other combinations or in isolation without departing from the scope of the disclosure.Brief description of the Drawings
[0013] FIG. 1 is a flowchart of a method according to various examples.
[0014] FIG. 2 schematically illustrates aspects with respect to an AHP hierarchy according to various examples.
[0015] FIG. 3 is a block diagram of a device according to various examples.Detailed description
[0016] Some examples of the present disclosure generally provide for a plurality of circuits or other electrical devices. All references to the circuits and other electrical devices and the functionality provided by each are not intended to be limited to encompassing only what is illustrated and described herein. While particular labels may be assigned to the various circuits or other electrical devices disclosed, such labels are not intended to limit the scope of operation for the circuits and the other electrical devices. Such circuits and other electrical devices may be combined with each other and / or separated in any manner based on the particular type of electrical implementation that is desired. It is recognized that any circuit or other electrical device disclosed herein may include any number of microcontrollers, a graphics processor unit (GPU) , integrated circuits, memory devices (e.g., FLASH, random access memory (RAM) , read only memory (ROM) , electrically programmable read only memory (EPROM) , electrically erasable programmable read only memory (EEPROM) , or other suitable variants thereof) , and software which co-act with one another to perform operation (s) disclosed herein. In addition, any one or more of the electrical devices may be configured to execute a program code that is embodied in a non-transitory computer readable medium programmed to perform any number of the functions as disclosed.
[0017] In the following, embodiments of the disclosure will be described in detail with reference to the accompanying drawings. It is to be understood that the following description of embodiments is not to be taken in a limiting sense. The scope of the disclosure is not intended to be limited by the embodiments described hereinafter or by the drawings, which are taken to be illustrative only.
[0018] The drawings are to be regarded as being schematic representations and elements illustrated in the drawings are not necessarily shown to scale. Rather, the various elements are represented such that their function and general purpose become apparent to a person skilled in the art. Any connection or coupling between functional blocks, devices, components, or other physical or functional units shown in the drawings or described herein may also be implemented by an indirect connection or coupling. A coupling between components may also be established over a wireless connection. Functional blocks may be implemented in hardware, firmware, software, or a combination thereof.
[0019] Hereinafter, techniques for determining a (final) set of search results based on a first weight associated with a first set of search results and a second weight associated with a second set of search results are disclosed. The first set of search results is obtained based on one or more keywords associated with a search query and the second set of search results is obtained based on one or more contextual vectors associated with the search query. The first weight and the second weight respectively represent how well the first set of search results and the second set of search results match the search query.
[0020] According to this disclosure, a search query may be a series of letters, a word, a phrase, or a sentence, e.g., in a natural language, that a user types to obtain search results that may match the search query. A search query may comprise a database query or a web search query.
[0021] For example, the search query may be associated with video information or audio information, and both the first set of search results and the second set of search results may respectively comprise at least one video file / document or at least one audio file / document.
[0022] In general, a keyword is also known as an index term, a subject term, a subject heading, or a descriptor, which is a term that captures essence of, e.g., a search query. Keywords associated with a search query may comprise a word, phrase, or alphanumerical term. They are created by analysing a search query either manually or automatically with automatic indexing or more sophisticated methods of keyword extraction. Keywords may either come from a controlled vocabulary or be freely assigned. The one or more keywords associated with a search query may be created by analysing a search query either manually or automatically with automatic indexing or more sophisticated methods of keyword extraction.
[0023] For example, the one or more keywords associated with a search query may be created or determined manually by identifying one or more main terms that convey a core meaning of the search query using terms from a controlled vocabulary. Additionally or optionally, synonyms and variations of the one or more keywords may be taken into account. Optionally or alternatively, the one or more keywords associated with a search query may be created or determined using automated methods such as natural language processing techniques and / or keyword extraction algorithms. The natural language processing techniques may involve analysing the search query to identify relevant words or phrases that capture the essential meaning of the query. The keyword extraction algorithms are used to identify and extract the most important words or phrases from a search query.
[0024] For example, Latent Semantic Analysis may be utilized to analyse a search query to find latent semantic structures and extract keywords that capture main themes. For the query "benefits of yoga for mental health" , the latent semantic analysis may highlight "benefits" , "yoga" , and "mental health" as keywords. As another example, machine learning based method, e.g., Bidirectional Encoder Representations from Transformers (BERT) , may be used to understand context and semantics of a search query, and extract relevant keywords. For the search query "how to bake a chocolate cake" , BERT may recognize "bake" , "chocolate" , and "cake" as keywords by understanding their contextual relevance. As a further example, Rapid Automatic Keyword Extraction (RAKE) as disclosed in non-patent literature –Rose S, Engel D, Cramer N, Cowley W. Automatic keyword extraction from individual documents. Text mining: applications and theory. 2010 Mar 26: 1-20. –may be used to determine one or more keywords from a search query. For the search query "machine learning for natural language processing" , RAKE might extract "machine learning" and "natural language processing" as key phrases.
[0025] Generally, a contextual vector is also known as a word embedding, which is a vector representation of a word. The contextual vector may be a real-valued vector that encodes the meaning of the word in such a way that the words that are closer in the vector space are similar in meaning. Contextual vectors can be obtained using language modelling and feature learning techniques, where words or phrases from a vocabulary are mapped to vectors of real numbers. Methods to generate this mapping may include neural networks, dimensionality reduction on the word co-occurrence matrix, probabilistic models, explainable knowledge base method, and explicit representation in terms of the context in which words appear.
[0026] For example, Word2Vec may be used to obtain, or determine, or generate one or more contextual vectors associated with a search query. Word2Vec comprises a group of shallow, two-layer neural network models that produce contextual vectors based on the context in which words appear. As another example, Transformer-based Sentence Embeddings, e.g., Sentence-BERT as disclosed in non-patent literature -Reimers N, Gurevych I. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv: 1908.10084.2019 Aug 27. -may be used to obtain, or determine, or generate one or more contextual vectors associated with a search query. As a further example, FastText, e.g., as disclosed in non-patent literature –Joulin A, Grave E, Bojanowski P, Douze M, Jégou H, Mikolov T. Fasttext. zip: Compressing text classification models. arXiv preprint arXiv: 1612.03651.2016 Dec 12. –may be used to obtain, or determine, or generate one or more contextual vectors associated with a search query, e.g., by summing or averaging the FastText vectors of words in the search query.
[0027] According to this disclosure, either the first set of search results or the second set of search results may comprise a single set or list of search results, e.g., video or audio documents / files, in which all the search results are ranked in a way that their relevancies decrease from the highest to the lowest. For example, the first set of search results may be ranked based on keyword relevance, i.e., how closely each of the search results of the first set matches the one or more keywords associated with the search query, while the second set of search results may be ranked based on contextual relevance, i.e., how well each of the search results of the second set fits or matches the one or more contextual vectors associated with the search query.
[0028] According to various examples, the first weight associated with the first set of search results and the second weight associated with the second set of search results may be determined based on a decision-making technique. For example, the first weight associated with the first set of search results and the second weight associated with the second set of search results may be determined based on an analytic hierarchy process (AHP) , which is also known as analytical hierarchy process.
[0029] The AHP provides a systematic approach to decision-making and helps in dealing with complex and multi-criteria decision problems. It has been widely used in various fields, including business, engineering, project management, and policy-making, to support decision-making processes. The core concept of the AHP is to systematically evaluate and prioritize alternatives by breaking down a decision problem into a hierarchy, making pairwise comparisons to determine relative importance, and aggregating the results to make informed decisions based on the aggregated results, e.g., the first weight and the second weight.
[0030] Alternatively, it is possible to use other decision-making techniques, e.g., analytic network process, Multi-Criteria Decision Analysis, Simple Additive Weighting (also known as weighted sum model or weighted linear combination) , the Technique for Order of Preference by Similarity to Ideal Solution, or a decision tree.
[0031] According to various examples, the (final) set of search results may be determined by ranking, using a weighted random algorithm or a stochastic universal sampling algorithm, each search result of the first set and the second set based on the first weight and the second weight.
[0032] FIG. 1 illustrates aspects with respect to a method 1000 for determining a set of search results based on a search query. The set of search results is determined based on a first weight associated with a first set of search results and a second weight associated with a second set of search results. The first set of search results is obtained based on one or more keywords associated with the search query and the second set of search results is obtained based on one or more contextual vectors associated with the search query. The first weight and the second weight respectively represent how well the first set of search results and the second set of search results match the search query.
[0033] The method 1000 can be executed by at least one processor upon loading program code. For example, the method 1000 could be executed by a processor of a server running a database management system, upon loading program code from a respective memory. Details with respect to such a method 1000 will be explained below.
[0034] Optional boxes are labeled with dashed lines.
[0035] Block 1100: obtaining a first set of search results based on one or more keywords associated with a search query.
[0036] Block 1200: obtaining a second set of search results based on one or more contextual vectors associated with the search query.
[0037] According to various examples, block 1100 and block 1200 may be performed together. For example, the first set of search results and the second set of search results may be stored in a database after being respectively obtained based on the one or more keywords and the one or more contextual vectors and the first set of search results and the second set of search results are linked or mapped to the search query. Thus, the first set of search results and the second set of search results can be obtained together, based on the query, from the database, e.g., in a way of postprocessing.
[0038] Alternatively or optionally, it is also possible to separately obtain the first set of search results and the second set of search results. I. e., block 1100 and block 1200 may be performed separately or in parallel, e.g., in a way of real-time processing. It is also possible to obtain either the first set of search results or the second set of search results from a database, while the other set of search results is obtained or retrieved from a further database based on the query.
[0039] Optionally or additionally, the method 1000 may comprise: block 1500, obtaining the search query. For example, the search query may be obtained from a user device, e.g., a wireless communication device, a desktop computer, a laptop computer, a tablet, or a device embedded in a car. For example, a user of the user device may input the search query via a web-based user interface or a user interface of a dedicated client software.
[0040] Optionally or additionally, i.e., block 1100, may comprise: block 1110, determining the one or more keywords based on the search query, and block 1120, obtaining the first set of search results based on the one or more keywords.
[0041] For example, the one or more keywords may be obtained using any one of Latent Semantic Analysis, a machine learning based method, and RAKE and the first set of search results may be obtained using an existing keyword search technique and based on the one or more keywords.
[0042] Optionally or additionally, i.e., block 1200, may comprise: block 1210, determining, using a machine-learning algorithm, the one or more contextual vectors based on the search query, and block 1220, obtaining the second set of search results based on the one or more contextual vectors.
[0043] For example, the one or more contextual vectors may be obtained using any algorithm of Word2Vec, Transformer-based Sentence Embeddings, and k-nearest neighbours and the second set of search results may be obtained using an existing contextual search technique and based on the one or more contextual vectors.
[0044] According to various examples, the first set of search results and the second set of search results may be obtained from the same database storing the same data or documents based one different searching techniques, i.e., keyword search and contextual search. It is also possible to obtain the first set of search results and the second set of search results respectively from two databases storing the same data or documents based on keyword search and contextual search.
[0045] Block 1300: determining a first weight associated with the first set of search results and a second weight associated with the second set of search results and the first weight and the second weight respectively represent how well the first set of search results and the second set of search results match the search query.
[0046] According to various examples, said determining of the first weight and the second weight, i.e., block 1300, may comprise: block 1310, determining a plurality of criteria for evaluating how well the first set of search results and the second set of search results match the search query, block 1320, determining, for each of the plurality of criteria, a respective value representing a level of importance associated with a corresponding criterion, block 1330, constructing a comparison matrix, each entry of the comparison matrix representing a relative importance associated with a respective pair of the plurality of criteria, block 1340, determining, based on the comparison matrix, a respective relative weight associated with each of the plurality of criteria, and block 1350, determining, based on the respective relative weight associated with each of the plurality of criteria, the first weight and the second weight.
[0047] FIG. 2 schematically illustrates aspects with respect to an AHP hierarchy 2000 according to various examples. The AHP hierarchy 2000 may comprise a goal 2100, a plurality of criteria 2210, 2220, and 2230, and several alternatives, i.e., the first set of search results 2310 and the second set of search results 2320.
[0048] According to various examples, the problem or goal may be a selection of a better or best search result. The plurality of criteria 2210, 2220, and 2230 may comprise at least two of a count of words in the search query, previous user behaviours, and a volume of data in a database where the first set and the second set are obtained. The plurality of criteria 2210, 2220, and 2230 may be used for evaluating how well the first set of search results and the second set of search results match the search query. Details related to the plurality of criteria are described in Table 1.
[0049] Table 1: criteria for better or best search result
[0050] For each of the plurality of criteria 2210, 2220, and 2230, a respective value representing a level / intensity of (relative) importance associated with a corresponding criterion may be determined. For example, the level / intensity of importance associated with a corresponding criterion may be determined or defined using a nine-point scale. Details related to the nine-point scale are described in Table 2.
[0051] Table 2: nine-point scale
[0052] A comparison matrix may be constructed based on pairwise comparisons of the plurality of criteria 2210, 2220, and 2230. Each entry of the comparison matrix represents a relative importance associated with a respective pair of the plurality of criteria 2210, 2220, and 2230. For example, a comparison matrix may be shown in table 3.
[0053] Table 3: comparison matrix
[0054] I.e., the comparison matrix shown in Table 3 is as follows:
[0055] A respective relative weight associated with each of the plurality of criteria 2210, 2220, and 2230 can be determined based on the comparison matrix A. For example, the respective relative weight associated with each of the plurality of criteria 2210, 2220, and 2230 can be determined according to the following two scenarios.
[0056] Scenario 1:
[0057] Said determining of the respective relative weight associated with each of the plurality of criteria 2210, 2220, and 2230, i.e., block 1340, may comprise: determining a respective sum for each column of the comparison matrix, normalizing the comparison matrix by dividing each entry of the comparison matrix by the respective sum of its column, determining a respective arithmetic mean for each row of the normalized comparison matrix. The respective arithmetic mean of each row of the normalized comparison matrix corresponds to the respective relative weight associated with each of the plurality of criteria.
[0058] Accordingly, the relative weight associated with criteria 2210, i.e., a count of words in the search query, is 0.258; the relative weight associated with criteria 2220, i.e., previous user behaviours, is 0.637; and the relative weight associated with criteria 2230, i.e., a volume of data in a database where the first set and the second set are obtained, is 0.105.
[0059] Scenario 2:
[0060] Said determining of the respective relative weight associated with each of the plurality of criteria 2210, 2220, and 2230, i.e., block 1340, may comprise: determining a respective geometric mean for each row of the comparison matrix, determining a sum of the respective geometric mean of each row of the comparison matrix, and determining the respective relative weight associated with each of the plurality of criteria by dividing the respective geometric mean of each row of the comparison matrix by the sum of the respective geometric mean of each row of the comparison matrix.
[0061] The same as scenario 1, the relative weight associated with criteria 2210, i.e., a count of words in the search query, is 0.258; the relative weight associated with criteria 2220, i.e., previous user behaviours, is 0.637; and the relative weight associated with criteria 2230, i.e., a volume of data in a database where the first set and the second set are obtained, is 0.105.
[0062] According to various examples, rating standards for both the first set of search results 2310 and the second set of search results 2320 may be defined according to the nine-point scale described in Table 2. Rating standards associated with the plurality of criteria 2210, 2220, and 2230 are respectively described in Tables 4, 5, and 6.
[0063] Table 4: Rating standards associated with criteria 2210
[0064] Table 5: Rating standards associated with criteria 2220
[0065] Table 6: Rating standards associated with criteria 2230
[0066] According to various examples, said determining of the first weight and the second weight, i.e., block 1350, may be further based on a weighted sum model. I. e., The first weight and the second weight may be determined based on the respective relative weight associated with each of the plurality of criteria and further based on a weighted sum model.
[0067] For example, a user inputs “company” as the search query, i.e., one word in the search query, selects, in the previous 10 times of search requests, a search result from the second set of search results 5 times, and also a search result from the first set of search results 5 times. The volume of data in a database where the first set and the second set are obtained is more than 20 gigabits. By referring to corresponding values described in Tables 4, 5, and 6, the respective scores associated with each of the criteria 2210, 2220, and 2230 are described in Table 7.
[0068] Table 7: respective scores associated with each of the criteria 2210, 2220, and 2230
[0069] Accordingly, the first weight is equal to 5 *0.258 + 3 *0.637 + 3 *0.105, i.e., 3.516, and the second weight is equal to 1 *0.258 + 3 *0.637 + 5 *0.105, i.e., 2.694.
[0070] As another example, a user inputs “working at Harman company” as the search query, i.e., four words in the search query, selects, in the previous 10 times of search requests, a search result from the second set of search results 8 times, and also a search result from the first set of search results 2 times. The volume of data in a database where the first set and the second set are obtained is more than 20 gigabits. By referring to corresponding values described in Tables 4, 5, and 6, the respective scores associated with each of the criteria 2210, 2220, and 2230 are described in Table 8.
[0071] Table 8: respective scores associated with each of the criteria 2210, 2220, and 2230
[0072] Accordingly, the first weight is equal to 3 *0.258 + 1 *0.637 + 3 *0.105, i.e., 1.726, and the second weight is equal to 3 *0.258 + 5 *0.637 + 5 *0.105, i.e., 4.484.
[0073] Block 1400: determining a set of search results based on the first weight and the second weight.
[0074] According to various examples, said determining of the set of search results, i.e., block 1400, may comprise: block 1410, determining a sum of the first weight and the second weight, and block 1420, determining whether either the first set or the second set is empty or not. If neither the first set nor the second set is empty, the method 1000 goes to block 1430, generating a random number between zero and the sum of the first weight and the second weight and block 1440, determining whether the random number is greater than the first weight. If the random number is greater than the first weight, the method 1000 goes to block 1460, appending a first search result of the second set to the set of search results and removing the first search result of the second set from the second set. If the random number is not greater than the first weight, the method 1000 goes to block 1450, appending a first search result of the first set to the set of search results and removing the first search result of the first set from the first set. After performing either block 1450 or block 1460, the method 1000 goes back to block 1420. If either the first set or the second set is empty, the method 1000 goes to block 1470, appending, to the set of search results, remaining search results in either the first set or the second set.
[0075] In other words, said determining of the set of search results, i.e., block 1400, may comprise: block 1410, determining a sum of the first weight and the second weight, iterating blocks 1430, 1440, 1450, as well as 1460 until all search results of either the first set or the second set have been appended to the set of search results, and block 1470.
[0076] FIG. 3 is a block diagram of a device 9000 according to various examples. The device 9000 may comprise at least one processor 9020, at least one memory 9030, and at least one input / output interface 9010. The at least one processor 9020 is configured to load program code from the at least one memory 9030 and execute the program code. Upon executing the program code, the at least one processor 9020 performs the method 1000. For example, the device 9000 may be a server running a database management system.
[0077] Summarizing, techniques have been disclosed that facilitate determining a (final) set of search results based on a first weight associated with a first set of search results and a second weight associated with a second set of search results are disclosed. By determining the (final) set of search results based on the first weight and the second weight, search results in a specific set, which has a greater weight, e.g., the second set of search results, will get more chances to be appended to the front of the (final) set of search results, i.e., being appended to the (final) set of search results earlier. When all the search results in the set, which has a greater weight, are appended to the (final) set of search results, the rest or remaining search results in the set, which has a smaller weight, will be appended to the end of the (final) set of search results. As such, the search quality associated with small databases or very short search queries can be improved. In addition, by generating a random number between zero and the sum of the first weight and the second weight for selecting a specific search result from either the first set or the second set, even the weight associated with a specific set of search results is very small, it still has a chance to append a search result of the specific set to the front of the final set of search results, and thereby the end user could see the diversity of the search result.
[0078] Although the disclosure has been shown and described with respect to certain preferred embodiments, equivalents and modifications will occur to others skilled in the art upon the reading and understanding of the specification. The present disclosure includes all such equivalents and modifications and is limited only by the scope of the appended claims.
Claims
1.A computer-implemented method (1000) , comprising:- obtaining (1100) a first set of search results based on one or more keywords associated with a search query;- obtaining (1200) a second set of search results based on one or more contextual vectors associated with the search query;- determining (1300) a first weight associated with the first set of search results and a second weight associated with the second set of search results, wherein the first weight and the second weight respectively represent how well the first set of search results and the second set of search results match the search query;- determining (1400) a set of search results based on the first weight and the second weight.2.The computer-implemented method (1000) of claim 1, wherein said determining of the first weight and the second weight is based on an analytic hierarchy process.3.The computer-implemented method (1000) of claim 1 or 2, wherein said determining of the set of search results comprises:- ranking, using a weighted random algorithm or a stochastic universal sampling algorithm, each search result of the first set and the second set based on the first weight and the second weight.4.The computer-implemented method (1000) of any one of preceding claims, wherein said obtaining (1100) of the first set of search results comprises:- obtaining (1500) the search query;- determining (1110) the one or more keywords based on the search query;- obtaining (1120) the first set of search results based on the one or more keywords.5.The computer-implemented method (1000) of any one of the preceding claims, wherein said obtaining (1200) of the second set of search results comprises:- obtaining (1500) the search query;- determining (1210) , using a machine-learning algorithm, the one or more contextual vectors based on the search query;- obtaining (1220) the second set of search results based on the one or more contextual vectors.6.The computer-implemented method (1000) of any one of the preceding claims, wherein said determining (1300) of the first weight and the second weight comprises:- determining (1310) a plurality of criteria for evaluating how well the first set of search results and the second set of search results match the search query;- determining (1320) , for each of the plurality of criteria, a respective value representing a level of importance associated with a corresponding criterion;- constructing (1330) a comparison matrix, each entry of the comparison matrix representing a relative importance associated with a respective pair of the plurality of criteria;- determining (1340) , based on the comparison matrix, a respective relative weight associated with each of the plurality of criteria;- determining (1350) , based on the respective relative weight associated with each of the plurality of criteria, the first weight and the second weight.7.The computer-implemented method (1000) of claim 6, wherein said determining (1340) of the respective relative weight associated with each of the plurality of criteria comprises:- determining a respective sum for each column of the comparison matrix;- normalizing the comparison matrix by dividing each entry of the comparison matrix by the respective sum of its column;- determining a respective arithmetic mean for each row of the normalized comparison matrix;wherein the respective arithmetic mean of each row of the normalized comparison matrix corresponds to the respective relative weight associated with each of the plurality of criteria.8.The computer-implemented method (1000) of claim 6, wherein said determining (1340) of the respective relative weight associated with each of the plurality of criteria comprises:- determining a respective geometric mean for each row of the comparison matrix;- determining a sum of the respective geometric mean of each row of the comparison matrix;- determining the respective relative weight associated with each of the plurality of criteria by dividing the respective geometric mean of each row of the comparison matrix by the sum of the respective geometric mean of each row of the comparison matrix.9.The computer-implemented method (1000) of any one of claims 6 to 8, wherein said determining (1300) of the first weight and the second weight is further based on a weighted sum model.10.The computer-implemented method (1000) of any one of claims 6 to 9, wherein the plurality of criteria comprises at least two of a count of words in the search query, previous user behaviours, and a volume of data in a database where the first set and the second set are obtained.11.The computer-implemented method (1000) of any one of preceding claims, wherein said determining (1400) of the set of search results comprises:- determining (1410) a sum of the first weight and the second weight;- iterating (1420) the following steps until all search results of either the first set or the second set have been appended to the set of search results:- generating (1430) a random number between zero and the sum of the first weight and the second weight;- if the random number is equal to or smaller than the first weight: appending (1450) a first search result of the first set to the set of search results and removing the first search result of the first set from the first set;- if the random number is greater than the first weight: appending (1460) a first search result of the second set to the set of search results and removing the first search result of the second set from the second set;- appending (1470) , to the set of search results, remaining search results in either the first set or the second set.12.The computer-implemented method (1000) of any one of preceding claims, wherein the search query is associated with video information or audio information, and both the first set of search results and the second set of search results respectively comprise at least one video file or at least one audio file.13.A processing device (9000) comprising a processor (9020) and a memory (9030) , the processing device (9000) being configured to:- obtain a first set of search results based on one or more keywords associated with a search query;- obtain a second set of search results based on one or more contextual vectors associated with the search query;- determine a first weight associated with the first set of search results and a second weight associated with the second set of search results, wherein the first weight and the second weight respectively represent how well the first set of search results and the second set of search results match the search query;- determine a set of search results by ranking each search result of the first set and the second set based on the first weight and the second weight.14.The processing device of claim 12, wherein the processing device is configured to perform the method of any one of claims 2 to 12.15.A computer program product or a computer program or a computer-readable storage medium including program code, the program code is executed by at least one processor to cause the at least one processor to perform the method of any one of claims 1 to 11.
Citation Information
Patent Citations
User interfaces for search systems using in-line contextual queries
US8108385B2