System and method for sorting electronic documents based on query token density

By calculating and utilizing query token density values ​​and density scores, search results are sorted or reordered, and the problem of difficult to measure the correlation of search results in the prior art is solved, and the relevance and user experience of search results are improved.

CN114175012BActive Publication Date: 2025-05-27RAVX INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080046123.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-23
Filing Date
2020-04-23
Publication Date
2025-05-27
Estimated Expiration
2040-04-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively measure the correlation of search results, making it difficult for users to find relevant information in multiple unrelated search results.

Method used

By extracting the query tokens in the search query, determine the query token hit set within each document, and calculate the query token density value and density score, sort or reorder the search results based on these metrics.

Benefits of technology

Improve the accuracy of the relevance measurement of search results, so that the most relevant search results can be presented to users faster and more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114175012B_ABST
    Figure CN114175012B_ABST
Patent Text Reader

Abstract

A system including a search engine, the search engine being configured to: determine search results based on a search query and a search query context, extract query tokens from the search query, determine (multiple) query token hit sets within each search result document, each query token hit set including (multiple) query token hits within a defined proximity range of a centrally positioned query token within a defined proximity range, determine for each query token hit set within each document a query token density value (QTDV) between each query token hit and the centrally positioned query token, each QTDV being based on a distance between each query token hit and the centrally positioned query token, determine for each query token hit set a query token density score (QTDS), determine for each document a document density score (DDS), rank or re-rank each document within the search results based on the DDS, and transmit the ranked / re-ranked search engine results page for presentation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This disclosure claims priority to U.S. Provisional Patent Application No. 62 / 837,428, filed on April 23, 2019, entitled "SYSTEMS AND METHODS FOR RANKING ELECTRONIC DOCUMENTS BASED ON QUERY TOKEN DOCUMENT DENSITIES", the entire content of which is incorporated herein by reference. Background Technical Field

[0003] This disclosure generally relates to the field of electronic document search. More specifically, the disclosed embodiments relate to computerized systems and methods for electronic document search that rank, re - rank, and / or measure search results by relevance. Background Art

[0005] Users typically utilize search engines to search for queries to obtain quick answers to questions. Unfortunately, users often need to sift through multiple irrelevant search results before the documents relevant to their queries are revealed. The use of boolean operators typically makes the problem more complex. If boolean operators are used in a search query, improper use of the boolean operators may inadvertently omit relevant documents from the overall search results. Thus, users may choose to enter natural language in the search query. However, conventional natural language search algorithms can have problems deciphering relatively long multi-word natural language queries, natural language queries including multiple concepts (e.g., related and / or distinct), natural language queries including hybrid search patterns (e.g., natural language search and entity search), and natural language queries including statistically common (e.g., domain-specific) terms in a corpus. Illustrative search queries demonstrating such problems include: "fraud by misappropriation", "motion to dismiss", "accepted as true", "2nd DCA", "private cause of action Constitutional right to privacy employer", "What statute requires a power of attorney to be recorded in conveyances act?", etc. Such problems can further "bury" relevant documents in multiple irrelevant search results, thus not presenting the user's answer quickly at all.

[0006] A related problem is how to adequately measure the relevance of search results. While this problem is crucial for measuring the objective utility of specific documents returned by a specific algorithm in response to a specific query, there are relatively few methods for computing search relevance, and each method is constrained by specific assumptions, advantages, and / or disadvantages. Thus, there is a great need for a reliable new method to improve the state of the art of search relevance measurement techniques.

[0007] Accordingly, there is a need for improved algorithms for sorting, re-sorting, and / or measuring search results based on search query search terms to improve the search engine result set. SUMMARY OF THE INVENTION

[0008] In a first aspect, a system for sorting electronic documents may include a search application device and a search engine device. The search application device includes a processor and a non-transitory computer-readable medium that includes program instructions. When executed by the processor, the program instructions cause the processor to: receive a search query and a search query context from a client device via one or more graphical user interfaces. The search engine device may include a processor and a non-transitory computer-readable medium that includes program instructions. When executed by the processor, the program instructions cause the processor to: determine search results based on the search query and the search query context, extract query tokens from the search query, determine one or more query token hit sets within each document of the search results, where each query token hit set includes one or more query token hits within a defined proximity range of a centrally located query token within the defined proximity range, determine a query token density value (QTDV) between each query token hit and the centrally located query token for each query token hit set within each document, where each QTDV is based on the distance between each query token hit and the centrally located query token of each query token hit, determine a query token density score (QTDS) for each query token hit set within each document, determine a document density score (DDS) for each document, sort or re-sort each document based on the DDS determined for each document within the search results; and transmit the sorted or re-sorted search engine results page to the search application device for presentation via the client device.

[0009] In a second aspect, a search engine may include a processor and a non-transitory computer-readable medium that includes program instructions, which when executed by the processor cause the processor to: execute a search query and a search query context to determine initial search results, extract query tokens from the search query, extract a defined proximity range, determine one or more query token hit sets within each document of the initial search results, where each query token hit set includes one or more query token hits within a defined proximity range of a centrally located query token within the defined proximity range, determine a query token density value (QTDV) between each query token hit and the centrally located query token for each query token hit set within each document, where each QTDV is based on the distance between each query token hit and the centrally located query token of each query token hit, determine a query token density score (QTDS) for each query token hit set within each document, determine a document density score (DDS) for each document, and re-sort each document based on the DDS determined for each document within the initial search results to generate a re-sorted search engine results page.

[0010] In a third aspect, a computer-implemented method for sorting electronic documents may include: receiving a search query and a search query context via a search application device, determining search results via a search engine based on the search query and the search query context; and performing a query token density algorithm of the search engine to: extract query tokens from the search query, determine one or more query token hit sets within each document of the search results, where each query token hit set includes one or more query token hits within a defined proximity range of a centrally located query token within the defined proximity range, determine a query token density value (QTDV) between each query token hit and the centrally located query token for each query token hit set within each document, where each QTDV is based on the distance between each query token hit and the centrally located query token of each query token hit, determine a query token density score (QTDS) for each query token hit set within each document, determine a document density score (DDS) for each document, sort or re-sort each document within the search results based on the DDS determined for each document, and transmit the sorted or re-sorted search engine results page to the search application device for presentation via a client device.

[0011] The additional features and advantages described herein will be set forth in the following detailed description, and will in part be apparent from the description, or may be learned by practice of the aspects described herein, including the following detailed description, the claims, and the drawings.

[0012] It should be understood that both the foregoing general description and the following detailed description describe aspects, and they are intended to provide an overview or framework for understanding the nature and characteristics of the claimed subject matter. The drawings are included to provide a further understanding of the aspects, and the drawings are incorporated into and constitute a part of this specification. The drawings illustrate the aspects described herein and, together with the specification, are used to explain the principles and operations of the claimed subject matter. Brief Description of the Drawings

[0013] The embodiments set forth in the drawings are illustrative and exemplary in nature and are not intended to limit the subject matter defined by the claims. The following detailed description of the illustrative embodiments can be understood when read in conjunction with the following drawings, in which like reference numerals indicate like structures, and in which:

[0014] Figure 1 Illustrative search results and an illustrative document of the search results are depicted in accordance with one or more embodiments shown and described herein, where multiple query tokens have been identified by multiple query token identifiers;

[0015] Figure 2Depicts an illustrative data file depicting multiple query tokens associated with each document of the trace and search results according to one or more embodiments shown and described herein;

[0016] Figure 3 Depicts a flowchart of an illustrative process for ranking or sorting electronic documents of search results based on document density scores according to one or more embodiments shown and described herein;

[0017] Figure 4 Depicts illustrative document density-based search results for ranking or re-ranking based on document density scores according to one or more embodiments shown and described herein;

[0018] Figure 5 Depicts an illustrative document for determining multiple QTDVs based on the distance between query tokens according to one or more embodiments shown and described herein; and

[0019] Figure 6 Depicts an illustrative QTD system according to one or more embodiments shown and described herein. Detailed Description

[0020] Aspects of determining the query token document density of each of a plurality of search result documents and ranking or re-ranking the search result documents based on the determined query token document density will now be described in detail, examples of which are illustrated in the accompanying drawings. Wherever possible, the same reference numerals will be used throughout the drawings to indicate the same or like parts.

[0021] Aspects of the present disclosure relate to systems and methods for searching a corpus of electronic documents. According to aspects, a corpus of electronic documents can include, but is not limited to, court opinions, statutes, secondary materials, news articles (e.g., Lexis, LexisUni, etc.), legal-related cases, or similar content from various countries (e.g., the United States, Canada, Australia, the United Kingdom, etc.). Aspects of the present disclosure extend to different content types, domains, and / or languages. According to various aspects, natural language processing can be used to search electronic documents such that a user can type query terms into a text field without the need for Boolean connectors. For example, a user can type "concerned about future identity theft" into a search box of a search engine interface without using any Boolean connectors. In other aspects, Boolean connectors can be used, but are ignored by the search algorithm.

[0022] In accordance with aspects of the present disclosure, one or more natural language search algorithms may search a document corpus for relevant electronic documents that contain at least some of the query terms in a query. In some aspects, an electronic document of the corpus having query terms that are close to each other (e.g., clustered together) may be considered more relevant to a user-entered query than an electronic document of the corpus having query terms that are relatively dispersed throughout the document (e.g., not clustered together).

[0023] Aspects of the present disclosure analyze the query token density of a document. According to aspects, as described more fully herein, query token density may be a weighted measure of distance between co-occurrences of query tokens (e.g., natural language search query terms) within a particular neighborhood (e.g., a defined proximity range) around a particular query token location (e.g., a portion of a document). In some aspects, query tokens that appear within the defined proximity may be given a weight proportional to their distance from a query token located at the center of the defined proximity range.

[0024] As described herein, aspects of the present disclosure may reorder the order of original or initial search results based on query token proximity. In such aspects, in a reordered list that may be presented to a user via an electronic display, the ranking of an electronic document having clustered query terms (e.g., query terms that are relatively close to each other) may be elevated to be higher than that of an electronic document having dispersed query terms (e.g., query terms that are relatively far from each other). In other aspects, no such reordering may occur. In such aspects, as described herein, query token proximity may have been considered during the determination or calculation of the initial search results.

[0025] In accordance with aspects of the present disclosure, a computer-implemented method may include determining query terms (e.g., query tokens, alternatively referred to herein) within each electronic document (e.g., each electronic document of search results) that are close to each other within a defined proximity range (e.g., query tokens within 5 tokens of each other, query tokens within 10 tokens of each other, etc.). According to aspects, the proximity range may be a predefined proximity range (e.g., preset prior to runtime) and / or a user-selectable proximity range (e.g., selectable specifically via a user interface, etc.).

[0026] Continuing with the example of the present text, a natural language query could be "worried about future identity theft". In such aspects, a query token identifier can be assigned to each query token (e.g., query term). For the sake of illustration, the query token "worried" can be identified by "A", the query token "future" can be identified by "B", the query token "identity" can be identified by "C", and the query token "theft" can be identified by "D". It should be understood that each corresponding query token identifier can include (multiple) letters, (multiple) numbers, (multiple) symbols, etc. to (e.g., uniquely) identify each query token of the search query. In this case, multiple query tokens can be combined to create a single query token (e.g., in the search query "worried about future identity theft", the query token "identity theft" can be interpreted as a single query token). According to various aspects, using multiple query token identifiers can simplify query token proximity analysis.

[0027] Figure 1 Depicts an illustrative search result and an illustrative document thereof according to various aspects of the present disclosure, in which multiple query tokens have been identified via multiple query token identifiers. Referring to Figure 1 , search result 100 (e.g., provided via a search engine results page on a user interface) can include document 1 ("Document 1") 102, document 2 ("Document 2") 112, document 3 ("Document 3") 122, etc. From Figure 1 view, in each document (e.g., document 1 102, document 2 112, document 3 122), each query token (e.g., "worried", "future", "identity", "theft", etc.) can be identified by its corresponding query token identifier (e.g., "A", "B", "C", "D", etc.). Additionally, in each document (e.g., document 1 102, document 2 112, document 3 122), for each query token (e.g., "worried", "future", "identity", "theft", etc.), other query tokens within a defined proximity range (e.g., of the search query) are identified.

[0028] From Figure 1 view, for example, in the first part of document 1 102, it has been determined that the first query token hit set 104 includes a query token hit "B" (e.g., "future") within the defined proximity range of the query token "C" (e.g., "identity"). Additionally, in the second part of document 1 102, it has been determined that the second query token hit set 106 includes a query token hit "A" (e.g., "worried") within the defined proximity range of the query token "D" (e.g., "theft"). Viewing the first query token hit set 104 with reference to the second query token hit set 106, the query token hit "B" is depicted as being relatively closer to the query token "C", and the query token hit "A" is depicted as being relatively farther from the query token "D".

[0029] Similarly, in view of Figure 1 , in the first part of Document 2 112, it has been determined that the third query token hit set 114 includes query token hits "A" and "B" within the defined proximity range of query token "D". Additionally, in the second part of Document 2 112, it has been determined that the fourth query token hit set 116 does not include (a) query token hit(s) (e.g., "A", "B", "C", "D", etc.) within the defined proximity range of query token "B".

[0030] Furthermore, in view of Figure 1 , in the first part of Document 3 122, it has been determined that the fifth query token hit set 124 includes query token hits "A", "B", and "A" is within the defined proximity range of query token "D". Additionally, in the second part of Document 3 122, it has been determined that the sixth query token hit set 126 includes query token hits "A", "C", and "A" is within the defined proximity range of query token "B". In view of the fifth query token hit set 124, the first instance of query token hit "B" and query token hit "A" are depicted as being relatively close to query token "D", while the second instance of query token hit "A" is depicted as being relatively further away from query token "D". Similarly, from the sixth query token hit set 126, query token hit "C" is depicted as being relatively close to query token "B", while the first instance of query token hit "A" and the second instance of query token hit "A" are depicted as being relatively further away (e.g., approximately equidistant) from query token "B".

[0031] Figure 2 An illustrative data file 202 (e.g., of database 637 as shown in Figure 6 ) depicting multiple query tokens associated with each document of the tracking and search results in accordance with aspects of the present disclosure is shown. Referring to Figure 2 , a first plurality of rows (e.g., row 204, row 206, etc.) may be associated with a first document (e.g., Document 1 102), a second plurality of rows (e.g., row 208, row 210, etc.) may be associated with a second document (e.g., Document 2 112), and a third plurality of rows (e.g., row 212, row 214, etc.) may be associated with a third document (e.g., Document 3 122).

[0032] From Figure 2 it can be seen that the data file 202 may record information associated with each query token hit set (e.g., query token hit sets 104, 106, 114, 116, 124, 126, etc.). Referring to Figure 2, a query token density score (QTDS) 216 can be associated with each line of the data file 202 (e.g., lines 204 - 214, etc.). In accordance with aspects of the present disclosure, each QTDS 216 can be determined based on the relative positions of the query tokens within each query token hit set (e.g., query token hit sets 104, 106, 114, 116, 124, 126, etc.). In aspects, a query token density value (QTDV) can be determined between query token pairs within each query token hit set, and the query token density values (QTDV) can be combined to calculate the QTDS 216. In some aspects, for example, the QTDV can be based on the distance (e.g., the number of tokens) between a query token in a query token hit set and another query token. In such aspects, for example, query tokens that are relatively closer to each other can receive a relatively higher QTDV than query tokens that are relatively farther from each other.

[0033] Brief reference Figure 1 , the document 1 102 can include a first query token hit set 104 (e.g., a query token hit "B" within a defined proximity of the query token "C") and a second query token hit set 106 (e.g., a query token hit "A" within a defined proximity of the query token "D"). Now turning to Figure 2 , reference Figure 1 , since the query token hit "B" is relatively closer to the query token "C", a relatively high QTDS 216 "3" in line 204 of the data file 202 is assigned to the first query token hit set 104, and since the query token hit "A" is relatively farther from the query token "D", a relatively low QTDS 216 "1" in line 206 of the data file 202 is assigned to the second query token hit set 106.

[0034] Similarly, document 2 112 may include a third query token hit set 114 (e.g., query token hits "A" and "B" within a defined proximity of query token "D") and a fourth query token hit set 116 (e.g., no query token hits within a defined proximity of query token "B"). Regarding the third query token hit set 114, since query token hit "A" is relatively further from query token "D", a relatively low QTDV may be assigned to it, and since query token hit "B" is relatively closer to query token "D", a relatively high QTDV may be assigned to it, such that QTDS 216 "3" in row 208 of data file 202 is assigned to the third query token hit set 114. Regarding the fourth query token hit set 116, since no query token is within the defined proximity of query token "B", QTDS 216 "0" in row 210 of data file 202 is assigned to the fourth query token hit set 116. According to aspects of the present disclosure, query token hit sets having (a) query token(s) outside the defined proximity range (e.g., the fourth query token hit set 116) are associated with a default QTDS "0".

[0035] Still further, document 3 122 may include a fifth query token hit set 124 (e.g., query token hits “A”, “B”, and “A” is within a defined proximity of query token “D”) and a sixth query token hit set 126 (e.g., query token hits “A”, “C”, and “A” is within a defined proximity of query token “B”). For the fifth query token hit set 124, since the first instance of query token hit “A” is relatively closer to query token “D”, a relatively high QTDV may be assigned to it, and since the second instance of query token hit “A” is relatively farther from query token “D”, a relatively low QTDV may be assigned to it. Additionally, since query token hit “B” is closer to query token “D” than the first instance of query token hit “A”, a QTDV initially higher than that of the first instance of query token hit “A” may be assigned to it. However, in accordance with aspects of the present disclosure, the QTDV assigned to repeated query token hits within a defined proximity (e.g., the first instance of query token hit “A” and the second instance of query token hit “A”) may be increased by a predetermined factor (e.g., repeated query token hits within a defined proximity may indicate a higher relevance of that portion of the document). Thus, the QTDV initially assigned to the first and / or second instance of query token hit “A” may be increased or raised above the QTDV initially assigned to query token hit “B”. In view of row 212 of data file 202, when combining the QTDV of the fifth query token hit set 124, QTDS 216 “12” is assigned to the fifth query token hit set 124. Similarly, for the sixth query token hit set 126, since the first instance of query token hit “A” and the second instance of query token hit “A” are similar in terms of proximity to query token “B”, QTDV may be assigned to them. Additionally, since query token hit “C” is closer to query token “B” than both the first and second instances of query token hit “A”, a QTDV initially higher than that of the first and second instances of query token hit “A” may be assigned to it. However, as described herein, the QTDV assigned to repeated query token hits within a defined proximity may be increased by a predetermined factor. Thus, the QTDV initially assigned to the first and / or second instance of query token hit “A” may be increased or raised above the QTDV initially assigned to query token hit “B”. In view of row 214 of data file 202, when combining the QTDV of the sixth query token hit set 126, QTDS 216 “10” is assigned to the sixth query token hit set 126.

[0036] Still referring to Figure 2, after each query token hit set of each electronic document (e.g., Document 1 1102, Document 2 112, Document 3 122, etc.) to which the QTDS 216 has been assigned, a document density score (DDS) 218 can be determined for each electronic document. According to aspects of the present disclosure, the DDS 218 can be determined based on one or more QTDS 216 of the respective query token hit sets assigned to the electronic document. In some aspects, for example, the DDS 118 can be the sum of the QTDS 216 associated with each electronic document. Continuing with this example, referring to Figure 2 , adding the QTDS of the first query token hit set 104 (e.g., "3" in row 204) and the QTDS of the second query token hit set 106 (e.g., "1" in row 206) gives a DDS 118 of "4" for Document 1 102. Similarly, adding the QTDS of the third query token hit set 114 (e.g., "3" in row 208) and the QTDS of the fourth query token hit set 116 (e.g., "0" in row 210) gives a DDS 118 of "3" for Document 2 112. Further still, adding the QTDS of the fifth query token hit set 124 (e.g., "12" in row 212) and the QTDS of the sixth query token hit set 126 (e.g., "10" in row 214) gives a DDS 118 of "22" for Document 3 122.

[0037] Figure 3 FIG. depicts a flowchart of an illustrative process for ranking or sorting electronic documents of a search result based on a document density score according to aspects of the present disclosure. At block 302, a search query (e.g., a search string including query tokens) and search query parameters (e.g., a defined proximity range, a maximum document ranking value to be re-ranked, etc.) can be received. At block 304, query tokens can be extracted from the search query. As described herein, a user can input a search query including natural language (e.g., with or without boolean connectors) into a search box of a search engine interface. In response, the search algorithm of the present disclosure can be based on a query token database file 617 (e.g., Figure 6)Extract query tokens from the search query. According to various aspects, the query token database file 617 can be customized for a subject and / or domain of interest (e.g., query tokens associated with court opinions, etc.). In some aspects, the query token database file 617 can extract query tokens based on a list of synonyms (e.g., extract synonyms of the query terms entered in the search query as query tokens). In other aspects, the query token database file 617 can extract query tokens based on a mathematical vector associated with the query tokens. At block 306, depicted as optional in the dashed box, each extracted query token can be assigned a token identifier. At block 308, a defined proximity range (e.g., preset or user-selected) can be used to determine a result set of electronic documents from a corpus of documents including one or more query token hit sets. At block 310, as described herein, QTDS can be assigned to each query token hit set. At block 312, as described herein, DDS can be determined for each electronic document in the result set. At block 314, the result set can be sorted or re-sorted based on DDS (e.g., if the result set was initially sorted based on another or different sorting algorithm). According to various aspects, the result set can be re-sorted by simple re-sorting (e.g., numerically), sequential re-sorting, intelligent re-sorting (e.g., sorting based on the best results from several processes (e.g., a learning-to-rank process where QTD results are integrated as features into a learning-to-rank feature matrix)), etc.

[0038] Sort or re-sort the search results

[0039] Figure 4 Illustrative document density-based search results 400 for sorting or re-sorting based on document density score (DDS) in accordance with aspects of the present disclosure are depicted. Refer to Figure 4 , document 3 122 (e.g., having a DDS of "22" in Figure 2 ) is sorted in the first position in the document density-based search results 400. Here, if the search engine has determined search results 100 ( Figure 1 and Figure 6 , e.g., using another sorting algorithm 628N), then document 3 122 is re-sorted from the third position to the first position in the document density-based search results 400. Further refer to Figure 4 , document 1 102 (e.g., having a DDS of "4" in Figure 2 ) is sorted in the second position in the document density-based search results 400. Here, if the search engine has determined search results 100, then document 1 102 is re-sorted from the first position to the second position in the document density-based search results 400. Yet further, in Figure 4 , document 2 112 (e.g., inFigure 2 The DDS with "3" in it is sorted in the third position in the document density-based search results 400. Here, if the search engine has determined the search results 100, the document 2 112 is re-sorted from the second position to the third position in the document density-based search results 400.

[0040] According to aspects of the present disclosure, the pseudocode for sorting or re-sorting documents within a search engine results page (SERP) may include:

[0041]

[0042]

[0043] According to some aspects of the present disclosure, a document boost (DB) can be calculated to boost the initial search results 100 based on an associated DDS (e.g., Figure 1 ) for the documents within it to achieve document density-based search results. In such aspects, the DDS calculated for each respective document of the initial search results can be multiplied by the search result ranking of each respective document in the initial search results. Continuing with this example, the document 1 102 with a search ranking of "1" and a DDS of "4" in the initial search results can be assigned a DB of "4" (e.g., 1×4 = 4). Similarly, the document 2 112 with a search ranking of "2" and a DDS of "3" in the initial search results can be assigned a DB of "6" (e.g., 2×3 = 6), and the document 3 122 with a search ranking of "3" and a DDS of "22" in the initial search results can be assigned a DB of "66" (e.g., 3×22 = 66). In such aspects, in the document density-based search results (not shown), the document 1 102 can be re-sorted from the first position to the third position, the document 2 112 can remain in the second position, and the document 3 122 can be re-sorted from the third position to the first position. Here, it should be understood that in the document density-based search results, based on the DDS of the document, the document can be similarly de-boosted rather than boosted.

[0044] As described herein, sorting or re-sorting the search result documents results in higher visibility of the most relevant search result documents to the user within the search interface. According to aspects of the present disclosure, a search result document with multiple query tokens clustered together is more relevant than a search result document with few or no query tokens clustered together. By placing the most relevant search result documents at the top of the list, the user can first see the most relevant documents.

[0045] Assign a query token density score (QTDS)

[0046] In accordance with aspects of the present disclosure, each QTDS associated with each query token hit set can be assigned based on the distance between a query token in the query token hit set and another query token. In one aspect, each token / word within each document being analyzed (e.g., search results) can be a distance increment.

[0047] Figure 5 Illustrative document 502 is depicted that assigns multiple QTDSs thereto based on the distance between query tokens in accordance with aspects of the present disclosure. Referring Figure 5 to, illustrative search result document 502 can include query token hit set A 504, query token hit set B 506, query token hit set C 508, query token hit set D 510, query token hit set E 512, and query token hit set F 514. Referring Figure 5 to, each query token hit set 504, 506, 508, 510, 512, 514, etc. can be located within different parts of document 502. In Figure 5 this, for illustrative purposes, query token hit set A 504 and query token hit set E 512 are magnified from document 502. Referring to query token hit set A 504 and query token hit set E 512, a defined proximity range 530 has been established (e.g., preselected or selected / input via the user interface of a search engine). More specifically, as Figure 5 shown, the token / word count or distance increment "10" has been established as the defined proximity range 530. In some aspects, all query token hit sets 504 - 514 of document 502 have the defined proximity range 530. In other aspects, one or more of the query token hit sets 504 - 514 can have different defined proximity ranges (e.g., based on query tokens of interest, etc.).

[0048] In accordance with aspects described herein, when assigning QTDV and / or QTDS associated with query token hit sets of a search result document, half-width (HW) can be utilized. In such aspects, HW can be equal to half of the defined proximity range (e.g., 10 / 2 = 5, HW = 5). HW can be used to establish multiple tokens / words before and after a query token located at the center of the defined proximity range of the query token hit set as discussed herein. In accordance with aspects, in the context of legal-related documents, HW = 5 is determined to be a reasonable half-width. For other subject areas, domains, and / or content types, the half-width can be determined similarly (e.g., 3 for news articles, 10 for academic materials, etc.).

[0049] Regarding query token hit set A 504, for a first query token 540 located at the center 532 of a defined proximity range 530, a first set of tokens / words 552 (e.g., numbered 1 - 5) is before the first query token 540, and a second set of tokens / words 554 (e.g., numbered 1 - 5) is after the first query token 540. Here, each token / word of the first set 552 can be numbered starting from a number that starts at half of the defined proximity range 530 (e.g., half-width, HW = 10 / 2 = 5), and decreases in numerical order from the first query token 540 towards the start of the defined proximity range 530, as Figure 5 depicted. Similarly, each token / word of the second set 554 can be numbered starting from a number that starts at half of the defined proximity range 530 (e.g., half-width, HW = 10 / 2 = 5), and decreases in numerical order from the first query token 540 towards the end of the defined proximity range 530, as Figure 5 depicted. As described herein, such numbering of each token / word within the defined proximity range 530 enables query tokens that are closer to the first query token 540 to be assigned a relatively larger QTDV, and query tokens that are farther from the first query token 540 to be assigned a relatively smaller QTDV. In this way, query tokens that are relatively closer to the first query token 540 are elevated over query tokens that are relatively farther from the first query token 540 (e.g., indicating that query token hit set A 504 may be more relevant to the user's search query).

[0050] Still referring to Figure 5 , query token hit set A 504 does not include query token hits within the defined proximity range 530 other than the first query token 540. Thus, a QTDS of "0" can be assigned to query token hit set A 504 (e.g., similar to the fourth query token hit set 116 as described herein). Since there are no query token hits other than the first query token 540, no QTDV is determined. In a similar manner, query token hit set B 506 does not include query token hits within the defined proximity range 530 other than the second query token 550, and query token hit set D 510 does not include query token hits within the defined proximity range 530 other than the fourth query token 570. Thus, query token hit set B 506 and query token hit set D 510 can also each be assigned a QTDS of "0", and no QTDV is determined.

[0051] Regarding query token hit set E512, for the fifth query token 580 located at the center 532 of the defined proximity range 530, the first set of tokens / words 582 (e.g., numbered 1 - 5) is before the fifth query token 580, and the second set of tokens / words 584 (e.g., numbered 1 - 5) is after the fifth query token 580. Here, each token / word in the first set 582 can be numbered starting from a number that begins at half of the defined proximity range 530 (e.g., half-width, HW = 10 / 2 = 5), and decreases in numerical order from the fifth query token 580 towards the start of the defined proximity range 530, as Figure 5 depicted. Similarly, each token / word in the second set 584 can be numbered starting from a number that begins at half of the defined proximity range 530 (e.g., half-width, HW = 10 / 2 = 5), and decreases in numerical order from the fifth query token 580 towards the end of the defined proximity range 530, as Figure 5 depicted. As described herein, such numbering of each token / word within the defined proximity range 530 enables query tokens located closer to the fifth query token 580 to be assigned a relatively larger QTDV, and query tokens located farther from the fifth query token 580 to be assigned a relatively smaller QTDV. In this way, query tokens relatively closer to the fifth query token 580 are elevated above query tokens relatively farther from the fifth query token 580 (e.g., indicating that the query token hit set E512 may be more relevant to the user's search query).

[0052] In this case, still referring to Figure 5 , the query token hit set E512 includes, within the defined proximity range 530, in addition to the fifth query token 580, a query token hit 586. Here, according to aspects of the present disclosure, a QTDV can be determined between query tokens in each query token hit set and combined to calculate a QTDS. Thus, a QTDV can be determined for the query token pair including the query token hit 586 and the fifth query token 580. Given Figure 5 , since the query token hit 586 is positioned "3" distance increments (e.g., 3 token / word counts) away from the fifth query token 580, the query token hit set E512 can be assigned a QTDS of "3". In a similar manner, the query token hit set F514 includes, within the defined proximity range 530, in addition to the sixth query token 590, a query token hit 596. Given Figure 5, since the query token hit 596 is located at a "4" distance increment (e.g., 4 tokens / word count) from the sixth query token 590, the query token hit set F514 can be assigned a QTDS of "2". Further, in a similar manner, the query token hit set C 508 includes, within the defined proximity range 530, query token hits 566, 567, and 568 in addition to the third query token 560. Given that Figure 5 , since the query token hit 566 is located at a "3" distance increment (e.g., 3 tokens / word count) from the third query token 560, the query token pair including the query token hit 566 and the third query token 560 is assigned a QTDV of "3". Similarly, since the query token hit 567 is located at a "1" distance increment (e.g., 1 token / word count) from the third query token 560, the query token pair including the query token hit 567 and the third query token 560 is assigned a QTDV of "5". Further, since the query token hit 568 is located at a "5" distance increment (e.g., 5 tokens / word count) from the third query token 560, the query token pair including the query token hit 568 and the third query token 560 is assigned a QTDV of "1". Thus, a QTDS of "9" (e.g., 3 + 5 + 1 = 9) can be assigned to the query token hit set C 508.

[0053] Reference Figure 5 , as described herein, the DDS of the search result document 502 can ultimately be determined to be 14 (e.g., 0 + 0 + 9 + 0 + 3 + 2 = 14) for the purpose of sorting or re - sorting the search result document 502 within document density - based search results similar to those described herein.

[0054] Further reference Figure 5 , in some aspects of the present disclosure, tokens at each position (e.g., 5, 4, 3, 2, 1 on each side of the centrally - located query token) may not be evaluated (e.g., some token positions may be skipped). In one aspect, for example, every other token position (e.g., within the defined proximity range) may be evaluated as described herein instead of each token position (e.g., within the defined proximity range). These aspects can reduce computation time and improve efficiency as performance - tuning enhancements (e.g., reducing computer resource consumption) while not compromising the relevance of the resulting sorted / re - sorted search results. This can support other efficiency aspects of the present disclosure (e.g., directly evaluating search result documents rather than requiring any pre - processing of search result documents, etc.).

[0055] Discount Function and Density Function

[0056] As described herein, query tokens that are located relatively closer to the center of the defined proximity range may be assigned a relatively large QTDV, and query tokens that are located relatively farther from the center of the defined proximity range may be assigned a relatively low QTDV. In this context, in accordance with aspects of the present disclosure, an increasingly large discount density value may be applied at each distance increment (e.g., token / word count) from the query token located at the center of the query token hit set.

[0057] In accordance with aspects, a discount function (DF) may be defined. In some aspects, the DF may be the reciprocal of the distance increment (e.g., token / word count) from the centrally located token (e.g., DF = 1 / HW). Here, continuing Figure 5 with the example, the DF may be 0.2 (e.g., 1 / 5 = 0.2). It should be understood that more complex DFs may be used (e.g., DF = 1 / log(HW), DF = 1 / (HW) 2 etc.).

[0058] Further in this aspect, a density function that applies the calculated DF (e.g., QTDV = 1 + (1 - distance to query token hit)(DF)) may be used to calculate the QTDV that may be assigned to a query token hit. Here, continuing in view of Figure 5 the example in, a query token hit at a distance increment of one from the first query token 540 may be assigned a QTDV of "1.8" (e.g., 1 + (1 - (1)(0.2)) = 1.8). Similarly, query token hits at distances of 2, 3, 4, and 5 distance increments from the first query token 540 may be assigned QTDVs of "1.6", "1.4", "1.2", and "1.0", respectively (e.g., towards the start and end of the defined proximity range 530). It should be understood that more complex density functions may be used (e.g., linear, probabilistic, comparative word-to-vector [word2vec] cosine, BERT, etc.).

[0059] Testing algorithms based on query term density (QTD)

[0060] Discounted Cumulative Gain (DCG) is a metric used to evaluate the quality of the search result ranking and / or the effectiveness of a search engine algorithm. Accordingly, the DCG of the QTD-based algorithms of the present disclosure is evaluated herein.

[0061] It is assumed that applying the QTD-based algorithm of the present disclosure to natural language search queries will result in improved human DCG (hDCG) (a relevance ranking determined by statistically combining ratings from human subject matter experts) and engaged DCG (eDCG) scores (a relevance ranking determined by statistically combining ratings from users / customers). Table 1 below details sample queries and their corresponding initial QTD hDCG results. In such aspects, three (3) independent subject matter experts rated a total of 40 queries at a ranking depth of twenty (20) during a blind test process.

[0062]

[0063] Table 1: Initial Selected QTD hDCG Results

[0064] Table 2 below details the corresponding initial QTD hDCG baseline test results. In such aspects, three (3) independent subject matter experts rated one thousand and eighty-six (1086) queries at a depth of ten (10) during a blind test process.

[0065]

[0066] Table 2: Initial QTD hDCG Baseline Test

[0067] QTD-Based Algorithm Adjustment

[0068] Aspects of the present disclosure may include adjusting QTDV (e.g., corresponding to a pair of query terms in a query token hit set) and / or QTDS (e.g., the combined QTDV of a query token hit set) as described herein. In such aspects, tokens / words (e.g., non-query tokens) within a defined proximity range but not in the search query may be analyzed.

[0069] In one aspect, QTDV and / or QTDS may be weighted higher or lower based on text (e.g., tokens) surrounding a particular query token (e.g., a centrally located query token, another token in a query token hit set). For example, if the text / token is within the defined proximity range, the QTDV and / or QTDS associated with that particular query token may be weighted higher or lower.

[0070] In another aspect, QTDV and / or QTDS may be weighted higher or lower based on high-value text or low-value text, respectively (e.g., a multiplier [e.g., 1.3, etc.] may be applied). According to aspects, high-value text may include, for example, citations (e.g., court opinions, regulations, etc.), entities found within a database (e.g., specific people, locations, etc.), links, Keycite TMFlags, semantic facts (e.g., the GDP of the United States in 2019, etc.). Here, for example, if the name of a specific person is a token within the defined proximity range, the weights of the QTDV and / or QTDS associated with the query token hit set may be higher or lower than those of the query token hit set that does not include the name of the specific person.

[0071] In yet another aspect, the QTDV and / or QTDS weights may be made higher or lower based on sentiment analysis terms and / or tokens that exhibit strong emotions or moods. In some aspects, if positive or negative sentiment terms (e.g., lawsuit, subpoena, etc.) or tokens that exhibit strong positive or negative emotions are within the defined proximity range, the QTDV and / or QTDS associated with the query token hit set may be respectively higher or lower in weight than those QTDV and / or QTDS that are not within the defined proximity range.

[0072] In a further aspect, the QTDV and / or QTDS weights may be made higher or lower based on other tokens within the defined proximity range (e.g., hidden markers in the document, specific items not in the search query, undesirable aspects of the document [e.g., source, content type, expiration date, etc.], etc.).

[0073] In yet a further aspect, the QTDV and / or QTDS weights may be made higher or lower based on document fields (e.g., the title section, annotation section, body section, etc. of the document). If the query token is within a specific document field, the QTDV and / or QTDS associated with the query token may be made higher or lower in weight.

[0074] In an even further aspect, when the query token appears in order in the search result document (e.g., in the search query order), when a synonym of the query token (e.g., dog or canine) is found within the defined proximity range, etc., the weights of the QTDV and / or QTD may be made higher.

[0075] Algorithm Expansion Based on QTD

[0076] According to various aspects of the present invention, the calculated QTDV (e.g., corresponding to a pair of query terms in the query token hit set), QTDS (e.g., the combined QTDV of the query token hit set), and / or DDS (e.g., the combined QTDS of the document) may be used to trigger other events (e.g., side display [e.g., side display in the interface], user interface (UI) changes to enhance document content [e.g., underlining, highlighting], etc.). In some aspects, one or more of the adjustments discussed herein (e.g., weights) may trigger another event. Accordingly, various combinations of the weights and expansions of the calculated QTDV, QTDS, and / or DDS are contemplated herein.

[0077] QTD as a search metric

[0078] According to a further aspect of the present invention, the calculated QTDV (e.g., corresponding to a pair of query terms in a query token hit set), QTDS (e.g., the combined QTDV of a query token hit set), and / or DDS (e.g., the combined QTDS of a document) can be used as a measure of the precision and / or relevance of a document, a search engine results page (SERP), a data set, etc. (e.g., QTD is used as a relevance yardstick rather than for ranking or re-ranking methods).

[0079] In some aspects, in its most basic form, the search results derived from the QTD algorithm as described herein can be evaluated against search results derived from another algorithm of the search engine's algorithm stack 626( Figure 6 ).

[0080] According to aspects, the DDS calculated for each of a plurality of documents in a SERP can be used as a raw score. In such aspects, the raw score can be used in a manner similar to relevance calculation methods, including term frequency-inverse document frequency (TF-IDF) and / or Okapi Best Match 25 (BM25).

[0081] In a further aspect, a maximum ranking value (e.g., up to 20 search results) can be selected for a SERP (e.g., via a search engine user interface, etc.). In such aspects, the DDS can be calculated for each document of the initial search results. Here, each search result document having a DDS that places it at or above the maximum ranking value can be assigned its corresponding ranking index to create a dual-ranked QTD value for each ranking (e.g., 1 - 20) in the SERP. In such aspects, the ranked DDS can be further evaluated. In some aspects, if the DDS monotonically decreases with ranking, the search results can exhibit ordered relevance and precision. In other aspects, if the DDS does not monotonically decrease with ranking (e.g., the ranking includes one or more anomalous DDSs), there may be a precision failure within the search results, misaligned documents, inaccessible documents, etc. at the anomalous ranking(s).

[0082] In a further aspect, DDS can be used to measure the overall relevance of a SERP. Specifically, a maximum ranking value (e.g., up to 20 search results) can be selected for the SERP (e.g., via a search engine user interface, etc.), the DDS can be calculated for each document of the initial search results, and each search result document having a DDS at or above the maximum ranking value can be assigned a rank (e.g., 1 - 20). Here, the DDS of the ranked search result documents can be summed to generate an overall SERP QTD score. Further in such aspects, various DDSs in the ranking can be evaluated (e.g., DDS at a specific rank, intermediate DDS rankings, etc.). For example, the DDS of each document at ranks 3, 5, 10, and 20 can be extracted to evaluate and compare the behavior at each rank (e.g., similar to the way DCG is calculated at various levels of the SERP to distinguish precision). In another example, intermediate DDS rankings (e.g., rank 3 and rank 4) can be compared with another relevance metric (e.g., DCG, precision, ERR, etc.) to validate QTD - based hypotheses (e.g., relatively high query - term density is proportional to relatively high search result document relevance) against computational assumptions inherent in other relevance - metric methods or to highlight anomalous results.

[0083] In an even further aspect, DDS can be used to limit recall. Here, as described herein, the DDS calculated for a single search result document, the DDS calculated into the maximum ranking value, and / or the DDS used to calculate the overall SERP QTD score can be used to determine a minimum precision threshold (e.g., QTD - based precision). In such aspects, search result documents having a DDS below the minimum precision threshold can be considered unimportant or irrelevant (e.g., especially when associated with low or unimportant TF / IDF scores and / or BM25 scores) and removed from the SERP (e.g., reducing the recall of the overall search query to only the more important results). For example, if the minimum precision threshold is "5", 5000 documents with a DDS of "3" can be removed to enable the system to present more relevant results relatively faster. In some aspects, the average rank at which precision drops (e.g., normalized for comparison in a manner similar to TF / IDF and / or BM25) can be determined to appropriately limit recall.

[0084] QTD Search Algorithm System

[0085] Figure 6 Illustrates an illustrative QTD system 600 for performing a QTD search algorithm to provide QTD - ranked documents in response to a search query and a search query context, according to aspects of the present disclosure. Refer to Figure 6, the QTD system 600 may include multiple client devices 602A - 602N, a search application device 604, a search engine device 606, and a network 608. Given that Figure 6 , the multiple client devices 602A - 602N, the search application device 604, and the search engine device 606 may be communicatively coupled via the network 608. Here, it should be understood that although Figure 6 depicts a specific number of components, the QTD system 600 may include any number of these components. Additionally, the functions provided by one or more components of the QTD system 600 may be combined into one component or distributed across multiple components. For example, a server farm including several primary servers and several backup servers may be used to implement the search application device 604 and / or the search engine device 606. Additionally, the search application device 604 and / or the search engine device 606 may be implemented by distributing various processing steps discussed herein across multiple servers. Similarly, the functions of each client device 602A - 602N, the search application device 604, and / or the search engine device 606 may be combined into a single device. It should be understood that the disclosed aspects may be implemented by computer devices or workstations organized as shown, computer devices or workstations organized in a distributed processing system architecture, or computer devices or workstations organized in countless suitable combinations of software, hardware, and / or firmware.

[0086] According to aspects of the present disclosure, the network 608 may be a shared, public, and / or private network, may encompass a wide area network or a local area network, and / or may be implemented through any suitable combination of wired and / or wireless communication networks. In some aspects, the network 608 may include an intranet or the Internet.

[0087] In accordance with aspects described herein, the search application device 604 may include a computer device (e.g., a network computer, a server, a mainframe, etc.). In some aspects, the search application device 604 may be configured as a dedicated computer (e.g., a specific machine) specifically designed to perform the functions described herein. For example, one or more of the disclosed processing steps may be implemented on a field programmable gate array (“FPGA”), an application specific integrated circuit (“ASIC”), or a suitable chipset. The search application device 604 may include one or more processors 610 that can be selectively activated and / or reconfigured by a computer program. In particular, the processor 610 may perform steps or methods consistent with the disclosed aspects by reading program instructions from the memory 612 and executing the program instructions. The memory 612 may include a non-transitory computer-readable medium that includes one or more memories (e.g., RAM, ROM, etc.) or storage devices (e.g., magnetic memory) that store data as well as computer programs / software. The memory 612 may store program modules that, when executed by the processor 610, perform one or more of the steps discussed herein. In accordance with aspects of the present disclosure, the search application device 604 may include a search application 614 (e.g., stored in the memory 612) configured to receive a search query and a search query context from a user. In particular, the search application 614 may be configured to provide one or more graphical user interfaces (GUIs) to receive user query input. In accordance with aspects, in view of Figure 6 , the search application 614 may generate a search query interface 616 and / or a search query context interface 618 for the user to interact with the search application 614. In some aspects, the search query interface 616 and the search query context interface 618 may be combined into one interface.

[0088] In accordance with aspects described herein, the search engine device 606 may include a computing device (e.g., a network computer, a server, a mainframe, etc.). In some aspects, the search engine device 606 may be configured as a dedicated computer (e.g., a specific machine) specifically designed to perform the functions described herein. For example, one or more of the disclosed processing steps may be implemented on a field-programmable gate array (“FPGA”), an application-specific integrated circuit (“ASIC”), or a suitable chipset. The search engine device 606 may include one or more processors 620 that may be selectively activated and / or reconfigured by a computer program. In particular, the processor 620 may perform steps or methods consistent with the disclosed aspects by reading program instructions from the memory 622 and executing the program instructions. The memory 622 may include non-transitory computer-readable media that include one or more memories (e.g., RAM, ROM, etc.) or storage devices (e.g., magnetic memory) that store data as well as computer programs / software. The memory 622 may store program modules that, when executed by the processor 620, perform one or more of the steps discussed herein. The search engine device 606 may further include a search engine 624, which, as described herein, is configured (e.g., processor, memory, program instructions, firmware, etc.) to process / perform (e.g., received from the search application device 604) search queries and search query contexts to generate a search engine results page (SERP) 636. Here, the generated SERP 636 may be stored in the search results database 638 (e.g., in a human-readable form for further processing) and / or directly sent to the search application device 604 (e.g., in a human-readable form for transmission to the client devices 602A - 602N). Here, although the search results database 638 is depicted as part of the search engine device 606, the search results database 638 may be located external to the search engine device 606 (e.g., accessible via the network 608) and / or may be part of the search engine 624. Additionally, although the search engine 624 is depicted as a separate component within the search engine device 606, the search engine 624 may be a computer program executable from the memory 622. In accordance with various aspects, the search engine 624 may include a search algorithm stack 626 that includes a plurality of executable search algorithms 628A - 628N (e.g., in parallel or sequentially, depending on the search engine 624 and the content construction) to identify search result documents from the (multiple) corpus databases 630. In particular, the search algorithm stack 626 may include a QTD search algorithm 628A to identify and sort / re-sort search result documents from the (multiple) corpus 630, as described herein.In accordance with various aspects, one or more of the plurality of search algorithms 628A - 628N may call additional parameters during execution (e.g., default settings, boosts to efficiency or performance, search engine - specific parameters, parameters used in extended functionality, parameters used when content or metadata is present or absent in the (multiple) corpus databases 630, etc.). In such aspects, the additional parameters may be read from a search - engine parameter database 637 (e.g., before and / or during execution of the respective search algorithms 628A - 628N). For example, with respect to the QTD search algorithm 628A of the present disclosure, such additional parameters and / or (multiple) QTD - specific parameters 639 (e.g., default QTD parameters, parameters received via the search - query context interface 618 as described herein, etc.) may be retrieved from the search - engine parameter database 637 and / or the search - query context interface 618 (e.g., to execute the QTD search algorithm 628A). For illustrative purposes, a user may input an emotion - analysis search query and may define / input a series of negative - emotion words as hidden terms to be used as part of the QTD query. Here, a flag indicating that the search query is intended to return negative - emotion results may be included in the search - query context interface 618 to trigger the use of the input hidden negative terms as QTD - specific parameters. According to some aspects, the search - algorithm stack 626 may further include a default - sorting algorithm 628B to initially sort the search - result documents from the (multiple) corpus databases 630, as described herein. For example, the search engine 624 may execute the default - sorting algorithm 628N to initially sort the search - result documents based on multiple different queries that the documents match (e.g., a document that matches more queries may be ranked higher than a document that matches fewer queries). In other aspects, the default - sorting algorithm 628N may initially sort the search - result documents based on document age, citation style, pointers to / from the document, etc.

[0089] Although the corpus databases 630 are depicted as external to the search - engine device 606, it should be understood that the search - engine device 606 may include the (multiple) corpus databases 630. In some aspects, the (multiple) corpus databases 630 may be housed on a separate computer device 631 accessible via the network 608. In accordance with various aspects, each corpus database 630 may include content (e.g., documents) and metadata associated with each respective content.

[0090] In view of Figure 6 , the search - application device 604 and the search - engine device 606 are depicted as separate components of the QTD system 600. However, in some aspects, the search - application device 604 and the search - engine device 606 may be combined into a single search device 632 (depicted as optional in the dashed box).

[0091] In accordance with various aspects described herein, the plurality of client devices 602A - 602N can include computer devices (e.g., personal computers, mobile phones, network computers, servers, mainframes, etc.). In some aspects, each of the plurality of client devices 602A - 602N can be configured as a dedicated computer (e.g., a specific machine) specifically designed to perform the functions described herein. For example, one or more of the disclosed processing steps can be implemented on a field programmable gate array (“FPGA”), an application specific integrated circuit (“ASIC”), or a suitable chipset. Each of the plurality of client devices 602A - 602N can include one or more processors 640A - 640N that can be selectively activated and / or reconfigured by a computer program. In particular, each processor 640A - 640N can perform steps or methods consistent with the disclosed aspects by reading program instructions from a memory 642A - 642N and executing the program instructions. The memory 642A - 642N can include non - transient computer - readable media, which includes one or more memories (e.g., RAM, ROM, etc.) or storage devices (e.g., magnetic memory) that store data as well as computer programs / software. The memory 642A - 642N can store program modules that, when executed by the processors 640A - 640N, perform one or more of the steps discussed herein.

[0092] In some aspects, the client device 602A may further include, for example, a web browser application 644A (e.g., stored in the memory 642A), and the web browser application 644A is configured to present a browser interface (e.g., a web-based interface) including one or more web pages on the display of the client device 602A. In particular, the web browser application 644A may present a web page including a search query interface 616 and / or a search query context interface 618 provided by the search application 614 of the search application device 604 (e.g., via the network 608). In such aspects, a user (e.g., a user of the client device 602A) may interact with the search application 614 via the search query interface 616 and / or the search query context interface 618 to perform a search query as described herein (e.g., by entering a search string into the search query interface 616 and / or entering search parameters into the search query context interface 618). In some aspects, the client device 602A may submit multiple search queries at once (e.g., via the search application 614 of the search application device 604). As described herein, the client device 602A may receive the SERP 636 from the search application device 604 (e.g., via the network 608) in response to the execution of a search query, and present the SERP 636 to the user via the browser interface. In other aspects, the client device 602A may further include a client-based search application 645A (e.g., similar to the search application 614 described herein) to directly interact (e.g., via the network 608) with the search engine device 606 to perform a search query and view search results, as described herein.

[0093] In other aspects, the client device 602N may further include an automatic search application 647N (e.g., a software test framework, a compiler script application, etc. stored in the memory 642N), and the automatic search application 647N is configured to directly interact (e.g., via an application programming interface, etc.) with the search application 614 of the search application device 604 to perform a search query (e.g., in an automatic manner) and receive search results, as described herein.

[0094] Still referring to Figure 6, the search query interface 616 of the search application device 604 may include various user interface elements for searching documents in the corpus database(s) 630. For example, the search query interface 616 may include a query text box (not shown) where a user can enter a search string including search terms, numbers, symbols, and / or other descriptors of content (e.g., tags) to be found via natural language (e.g., without Boolean logic operators) or structured language (e.g., search terms connected via Boolean logic operators). The search query interface 616 may further include a search button (not shown) for submitting the search query to the search engine 624 (e.g., via the search application device 604). The search query interface 616 may further display the SERP(s) 636 sorted as described herein that match the user's search query.

[0095] Further in view of Figure 6 , the search query context interface 618 of the search application device 604 may include various user interface elements for setting parameters associated with the search query. For example, the search query context interface 618 may include a drop-down box (not shown), filters, etc., where the user can select parameters associated with the search query at the drop-down box (not shown), filter program. Optional parameters may include a defined proximity range, a maximum document sort value to be sorted or re-sorted, a description of one or more corpora to be searched (e.g., in the corpus database(s) 630 or other specified databases), a default search algorithm of the search algorithm stack 626 to be executed (e.g., and associated parameters), a dedicated and / or non-default search algorithm of the search algorithm stack 626 to be executed (e.g., and associated parameters) (e.g., the QTD search algorithm 628A as described herein), metadata associated with the search (e.g., date, user identification, content restrictions or bans, extended queries, query intent information, etc.), etc.

[0096] QTD Search Metric System

[0097] Still referring to Figure 6, in accordance with aspects of the present disclosure, the QTD system 600 can also be used as a QTD search metric system to provide relevance metrics for search result documents, SERP 636, etc., as described herein. In such aspects, a user (e.g., via the client device 602A) and / or an automated search application 647N (e.g., via the client device 602N) can interact with the search application 614 of the search application device 604 to simultaneously perform a grouping or listing of search queries. In one example, the automated search application 647N can include a search test application where QTD is used as a metric to display the relevance results of each search document result in the SERP 636. In another example, the automated search application 647N can include a compiler script that uses a group or list of search queries for regression testing or comparative testing.

[0098] According to aspects, the search query interface 616 of the search application 614 can include various user interface elements for entering a grouping or listing of search queries. For example, the search query interface 616 can include multiple text boxes (not shown) for entering multiple search queries in each grouping or each listing of search queries, one or more relatively large text boxes for cutting and pasting multiple search queries in each grouping or listing of search queries, an upload button for uploading multiple search queries in each grouping or each listing of search queries, etc. Further in such aspects, the search query context interface 618 of the search application 614 can include various user interface elements for setting parameters associated with each grouping or each listing of search queries (e.g., such that each query in each corresponding grouping or listing of search queries is processed identically). For example, the search query context interface 618 can include (multiple) text boxes (not shown) for entering a query identifier and / or context metadata associated with each grouping or each listing of search queries, as well as drop-down boxes (not shown), filters, etc., where the user can select parameters associated with each grouping or each listing of search queries. In some aspects, the parameters associated with a grouping or listing of search queries can be the same as those of another grouping or listing of search queries. In other aspects, the parameters associated with a grouping or listing of search queries are specific to that grouping or listing (e.g., an identifier or name of the query set, the (multiple) source of the query set, a description of the query set, etc.). In some aspects, a set of parameters prioritized by application order or priority can be a QTD-specific (multiple) parameter set 639.

[0099] According to various aspects, as a QTD search metric system, each group or each list of search queries may be submitted in a sequential manner (e.g., via the search application device 604) via the search query interface 616 and the search query context interface 618 for execution by the search engine 624 of the search engine device 606 as described herein. In such aspects, the SERP corresponding to each search query of each group or each list of search queries may be saved in the search result database 638 for subsequent metric analysis as described herein.

[0100] It should now be understood that the systems and methods described herein rank or re-rank electronic documents of search results by analyzing the proximity of query tokens to each other within each electronic document. More specifically, the systems and methods described herein rank or re-rank search result documents based on the query tokens of the search string (e.g., a natural language search string) and the proximity of each query token within a defined proximity range to other query tokens to generate an improved search engine result set. In the search result list presented to the user, the more relevant electronic documents are presented first.

[0101] Although specific embodiments have been shown and described herein, it should be understood that other changes and modifications may be made without departing from the spirit and scope of the claimed subject matter. Additionally, although various aspects of the claimed subject matter have been described herein, these aspects need not be utilized in a combined manner. Accordingly, the appended claims are intended to cover all such changes and modifications within the scope of the claimed subject matter.

Claims

1. A system for sorting electronic documents, the system comprising: A search application device, comprising: A processor and a non-transitory computer-readable medium, the non-transitory computer-readable medium including program instructions that, when executed by the processor, cause the processor to: Receive a search query and a search query context from a client device via one or more graphical user interfaces; and A search engine device, comprising: A processor and a non-transitory computer-readable medium, the non-transitory computer-readable medium including program instructions that, when executed by the processor, cause the processor to: Determine search results based on the search query and the search query context; Extract query tokens from the search query; Determine one or more query token hit sets within each document of the search results, where each query token hit set includes one or more query token hits within a defined proximity range of a centrally located query token within the defined proximity range; Determine a query token density value QTDV between each query token hit and the centrally located query token for each query token hit set within each document, where each QTDV is based on the distance between each query token hit and the centrally located query token of each query token hit; Determine a query token density score QTDS for each query token hit set within each document by combining the QTDV; Determine a document density score DDS for each document based on one or more QTDS assigned to the respective query token hit sets of the electronic document; Sort or re-sort each document based on the DDS determined for each document within the search results; and Transmit the sorted or re-sorted search engine results page to the search application device for presentation via the client device.

2. The system according to claim 1, wherein, The distance is based on the count of words or tokens between each query token hit and the centrally located query token of each query token hit.

3. The system according to claim 1, wherein, The query token hit set includes multiple query token hits within the defined proximity range of the centrally located query token, and wherein when executed by the processor of the search engine device, the program instructions further cause the processor to: Assign a relatively higher QTDV to a query token hit that is relatively closer to the centrally located query token of the query token hit set; And Assign a relatively lower QTDV to a query token hit that is relatively farther from the centrally located query token of the query token hit set.

4. The system according to claim 3, wherein, Each distance increment from the centrally located query token is associated with an increasingly discounted density value.

5. The system according to claim 3, wherein, Each QTDV is based on a density function that applies a discount function to each distance increment from the centrally located query token.

6. The system according to claim 5, wherein, The discount function is one of the following: 1 / HW, 1 / (log HW), or 1 / (HW 2 , where HW is the half-width equal to half of the defined proximity range.

7. The system according to claim 1, wherein, each distance increment from the query token centeredly located is associated with a reduced weight value.

8. The system according to claim 7, wherein, the query token hit set includes one query token hit within a defined proximity range of the query token centeredly located, and wherein when executed by the processor of the search engine device, the program instructions further cause the processor to: assign a QTDV to the query token hit set based on the reduced weight value.

9. The system according to claim 1, wherein, the query token hit set does not include a query token hit within a defined proximity range of the query token centeredly located, and wherein when executed by the processor of the search engine device, the program instructions further cause the processor to: assign a QTDV equal to zero to the query token hit set.

10. The system according to claim 1, wherein, when executed by the processor of the search engine device, the program instructions further cause the processor to: adjust at least one determined QTDV based on one or more of the following: duplicate query tokens within a defined proximity range, tokens including references, entities, links, flags, semantic facts or sentiment terms within a defined proximity range; query tokens within a specific document field; and query tokens in the same order in the document as in the search query.

11. The system according to claim 1, wherein, sorting or re - sorting each document in the search results further includes boosting or de - boosting each document based on its respective ranking in the initial search results.

12. The system according to claim 1, wherein, when executed by the processor of the search engine device, the program instructions further cause the processor to: use the DDS determined for each document as a relevance metric to sort or re - sort the search results relative to a maximum ranking value to measure the overall relevance of the sorted or re - sorted search engine results page, or remove a document from the sorted or re - sorted search engine results page.

13. The system according to claim 1, wherein, when executed by the processor of the search engine device, the program instructions further cause the processor to: receive, via a search query context interface, at least one of a defined proximity range, a corpus database to be searched, a maximum document ranking value for sorting or re - sorting, or a search algorithm to be applied.

14. A search engine, comprising: a processor and a non - transient computer - readable medium, the non - transient computer - readable medium including program instructions which, when executed by the processor, cause the processor to: execute a search query and a search query context to determine an initial search result; extract query tokens from the search query; extract a defined proximity range; Determine one or more query token hit sets within each document of the initial search results, where each query token hit set includes one or more query token hits within a defined proximity range of a centrally located query token within the defined proximity range; Determine a query token density value QTDV between each query token hit and the centrally located query token for each query token hit set within each document, where each QTDV is based on the distance between each query token hit and the centrally located query token of each query token hit; Determine a query token density score QTDS for each query token hit set within each document by combining the QTDVs; Determine a document density score DDS for each document based on one or more QTDSs assigned to the respective query token hit sets of the document; and Re - rank each document based on the DDS determined for each document within the initial search results to generate a re - ranked search engine results page.

15. The search engine according to claim 14, wherein, When executed by the processor, the program instructions further cause the processor to: Transmit the re - ranked search engine results page to at least one of a search application device or a search engine results database.

16. The search engine according to claim 14, wherein, The distance is based on the count of words or tokens between each query token hit and the centrally located query token of each query token hit.

17. The search engine according to claim 14: where when a query token hit set: Includes multiple query token hits within the defined proximity range of the centrally located query token, the program instructions further cause the processor to: Assign a relatively higher QTDV to a query token hit that is relatively closer to the centrally located query token of the query token hit set; and Assign a relatively lower QTDV to a query token hit that is relatively farther from the centrally located query token of the query token hit set; Includes one query token hit within the defined proximity range of the centrally located query token, the program instructions further cause the processor to: Assign a QTDV to the query token hit set based on a weight value that decreases with each distance increment from the centrally located query token; and Does not include a query token hit within the defined proximity range of the centrally located query token, the program instructions further cause the processor to: Assign a QTDV equal to zero to the query token hit set.

18. The search engine according to claim 14, further comprising: A search engine parameter database, where the defined proximity range is extracted from the search engine parameter database.

19. A computer - implemented method for sorting electronic documents, the method comprising: Receiving a search query and a search query context via a search application device; Determining search results via a search engine based on the search query and the search query context; and Execute the query token density algorithm of the search engine to: Extract query tokens from the search query; Determine one or more query token hit sets within each document of the search results, where each query token hit set includes one or more query token hits within a defined proximity range of a centrally located query token within the defined proximity range; Determine a query token density value QTDV between each query token hit and the centrally located query token for each query token hit set within each document, where each QTDV is based on the distance between each query token hit and the centrally located query token of each query token hit; Determine a query token density score QTDS for each query token hit set within each document by combining the QTDV; Determine a document density score DDS for each document based on one or more QTDS assigned to the respective query token hit sets of the electronic document; Sort or re - sort each document within the search results based on the DDS determined for each document; And Transmit the sorted or re - sorted search engine results page to the search application device for presentation via the client device.

20. The computer - implemented method according to claim 19, further comprising: When a query token hit set: Includes multiple query token hits within the defined proximity range of the centrally located query token: Assign a relatively high QTDV to the query token hits that are relatively closer to the centrally located query token of the query token hit set; And Assign a relatively low QTDV to the query token hits that are relatively farther from the centrally located query token of the query token hit set; When the query token hit set includes one query token hit within the defined proximity range of the centrally located query token: Assign the QTDV to the query token hit set based on a weight value that decreases with each distance increment from the centrally located query token; or When the query token hit set does not include a query token hit within the defined proximity range of the centrally located query token: Assign a QTDV equal to zero to the query token hit set.

Citation Information

Patent Citations

  • Retrieval method based on position characteristics

    CN106095780A

  • Target document obtaining method and application server

    CN108427702A