Classifying search result documents
By employing query-dependent and query-independent measures based on past interactions, the patent addresses the challenge of ranking access-restricted documents, improving search result relevance and reducing user input and network traffic.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2017-09-26
- Publication Date
- 2026-05-07
AI Technical Summary
Existing search technologies struggle to effectively rank access-restricted documents, such as personal emails, based on user interactions, as there is often insufficient interaction data available for these documents, rendering traditional user interaction data-based ranking techniques ineffective.
Utilize query-dependent and query-independent measures based on past interactions of other documents with similar features to determine presentation characteristics for access-restricted documents, leveraging bipartite graphs to model interactions between queries and documents, and their features, to rank and present search results.
Enhances the relevance of search results by prioritizing more relevant access-restricted documents, reducing user input and network traffic, and conserving computing resources by providing more accurate initial search results.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
background
[0001] Search engines provide information about various documents, such as web pages, images, text documents, multimedia content, and / or electronic communications. For example, in response to a search query, a search engine identifies one or more documents relevant to the query. The search engine ranks the documents based on their relevance to the query and / or other ranking signals, and delivers corresponding search results. The search results may include aspects of and / or links to the documents and may be delivered based on these rankings.
[0002] US 2007 / 0 150 470 A1 reveals a set of complementary techniques that significantly improve search and navigation results in businesses. At the heart of the US 2007 / 0 150 470 A1 revelation is a subject-matter or knowledge index called UseRank, which tracks the behavior of website visitors. This index focuses on four key characteristics of businesses: subject-matter expertise, work patterns, content recency, and group know-how. US 2007 / 0 150 470 A1 delivers useful, timely, cross-application, subject-matter-based search and navigation results. In contrast, traditional information-gathering technologies such as inverted indexes, NLP, or taxonomy address the same problem using a set of attributes that conflict with business needs: content population, word patterns, content presence, and statistical trends.Overall, the US 2007 / 0 150 470 A1 Baynote Search includes - an improvement over existing IR searches -, Baynote Guide - a set of community-driven navigations - and Baynote Insights - aggregated views of visitor interests and trends as well as content gaps.
[0003] AU 2008 261 113 A1 discloses a method for providing access for a user to selected objects from a plurality of target objects and sets of target object attributes that can be accessed via an electronic storage medium, wherein the users are connected via user terminals and data communication links to a target server system comprising the electronic storage medium, wherein the method comprises the following steps: automatically generating target profiles for target objects and sets of target object attributes stored in the electronic storage medium, wherein each of the target profiles is generated from the contents of an associated target object and set of target object attributes;Automatically generating at least one summary of the interests of a user target profile for a user at a user terminal, wherein each summary of the interests of a user target profile is generated from the target objects and sets of target object attributes accessed by the user; and enabling users to access the multitude of target objects and sets of target object attributes stored on the electronic storage medium via the target profiles.
[0004] WO 2007 / 124430 A2 provides a computer-implemented method that includes presenting a user (30) with a series of personalization levels for search results, including a personalized level, a global level that is not personalized, and a community level between the personalized and global levels. The user (30) receives a cue (550) indicating a desired level and a search query (52) consisting of one or more search terms. In response to the search query, a search result list (54) is generated. At least a portion of the search result list (54) is evaluated, at least partially, in response to the cue (550), and at least a portion of the evaluated search result list (54) is presented to the user (30). Further embodiments are also described.
[0005] US 2003 / 0123443A1 discloses a search engine that uses both record-based data and user activity data to develop, update, and refine ranking protocols and to identify words and phrases that lead to search ambiguities, so that the search engine can interact with the user to better respond to user queries and improve data collection from databases, intranets, and the Internet.
[0006] US 6,732,088 B1 discloses a method and system for facilitating searching in a data repository, such as the World Wide Web, by leveraging the collective ability of all users to submit search queries to the repository. First, a node-connection diagram is created of all search queries sent to the repository within a specific time period. In the case of the World Wide Web, the search queries would be directed to a specific search engine. In the diagram, each node represents a search query. A connection is established between two nodes when the two search queries are deemed related. A key consideration is that the determination of relatedness depends on the documents returned by the search queries, rather than on the actual terms in the search queries themselves.One criterion for relatedness could be, for example, that the two lists of the ten best documents returned for each search query have at least one document in common. US 6,732,088 B1 also describes a way to allow the user to browse the network of related search queries in an ordered manner: by following a path from a first cousin to a second cousin to a third cousin, and so on, until a certain number of results are reached. In this way, the user is shown related documents that were not returned as a result of their search query. A second key idea is that creating the search query graph transforms the use of the data collection by a single user (e.g., searching) into collaborative use: all users can access the knowledge base of search queries submitted by others.This is because each of the related search queries represents the knowledge of the user who submitted the search query. Summary
[0007] The claimed subject matter is defined in the accompanying claims. This patent specification is directed to technical features relating to using one or more document features of a given document in response to a query, and optionally one or more query features of the query, to determine a presentation characteristic for presenting a search result corresponding to the given document—and to providing the search result for presentation with the presentation characteristic in response to the query. In some implementations, the given document in response to the query may be an access-restricted document, such as an access-restricted document that only the user who submitted the query and, optionally, other users designated by that user can access.
[0008] In some implementations, measures associated with one or more document attributes and / or one or more query attributes can be used to determine the presentation characteristic. These measures can be based on past interactions of relevant users with other documents that share one or more of the document attributes with the given document, where several of the other documents are each distinct from the given document (and optionally distinct from each other). Using such measures makes it possible to leverage past interactions with other documents when determining interaction-based relevance of the given document, optionally without reference to any query-based past interactions specifically directed at the given document.In some implementations, the other documents contain documents that are themselves access-restricted, or access is restricted to them. This helps maintain security during the search process.
[0009] The presentation characteristic can be, for example, a presentation order or other display prominence. In various implementations, providing search results with the presentation characteristic offers the search results in a way that is more useful to the user compared to providing the search results without the presentation characteristic. The improvement is more objective than subjective, since the highest-ranked or most prominent search result is, on average, likely to be a more relevant result. This means that users are likely to spend less time accessing or examining other, less relevant search results to determine whether they are actually of interest. Consequently, the number of user inputs / selections received in relation to the search results is reduced on average.This reduction in user input / selections, while not necessarily achieved in every case when search results are presented with appropriate presentation characteristics, offers a significant reduction in network traffic for the network as a whole when scaled across many users of a search engine. Additionally, reducing further input / selections can, on average, reduce the use of computing resources within each user device. Such computing resources can include processing power resulting from handling user input and communicating with the network, as well as the power consumption resulting from screen power-on time.Reducing the screen's power-on time can result in users being more likely to end their search for specific information immediately after receiving the search results presented with the appropriate layout, without having to access and examine search results that are actually irrelevant to them. Screen usage, especially on mobile devices, can account for a significant portion of a user's device's power consumption, and therefore, reducing the screen's power-on time can be particularly advantageous.
[0010] In some implementations, when determining the presentation characteristic of a search result corresponding to a given document in response to a query, a query-dependent measure is generated for the given document and used to determine the presentation characteristic. In some of these implementations, the query-dependent measure is used to determine a score for the given document, and this score is used to rank the given document relative to other response documents for the query (e.g., based on their corresponding scores, which may also be based on corresponding query-dependent measures). For example, the query-dependent measure can be used to assign an initial score to the given document (e.g.,A score (based on the degree of match between the query and the given document) can be modified, and the modified score can be used to rank the given document relative to other response documents for the query. This ranking can be used, for example, to determine which response documents are initially used to provide relevant search results for presentation in response to the query, and / or to establish a presentation order (or other display prominence) for the search results.
[0011] In some implementations, the query-dependent measure for a given document in response to a query is determined based on measures of past interactions between query attributes and document attributes of the given document. Each measure can be based on a set of past interactions by relevant users with other documents exhibiting one or more of the document attributes, when those other documents were presented in response to relevant queries with one or more of the query attributes. Various past interactions can be used to determine the measures, such as selections of search results corresponding to the other documents in response to the relevant queries (e.g., a click-to-view fraction), document access counts, cursor tracking, and / or touch gestures.In some implementations, the other documents themselves may contain or be restricted to multiple access-restricted documents, such as inaccessible documents, each of which is personal to a corresponding other user and which the user cannot access.
[0012] In some implementations, a query-independent measure is generated for the given document and used additionally or alternatively to determine its presentation characteristics. In some of these implementations, the query-independent measure is based on measures of past interactions by relevant users with other documents that exhibit one or more of the document characteristics of the given document, when the other documents were presented in response to relevant requests that did not include any of the request characteristics. Accordingly, the query-independent measure can provide an indication of the general popularity of documents with the one or more document characteristics, while the request-dependent measure provides an indication of the popularity of documents with the one or more document characteristics in response to requests that do include the request characteristics.
[0013] In some implementations, a procedure is provided that includes receiving a request entered by a user via a user interface input device on the user's computing device. The procedure further includes identifying response documents in reply to the request, including an email sent to the user's email address. The procedure also includes identifying one or more document attributes for the email. The document attributes include at least one email attribute based on at least one of the following: sender content, based on its presence in a sender field of the email, and subject content, based on its presence in a subject field of the email.The method further comprises identifying one or more query characteristics for the query and generating a query-dependent measure for the email based on measures of past interactions between the query characteristics and the document characteristics, wherein each set of measures is based on a set of past interactions by relevant users with other documents exhibiting one or more of the document characteristics, when the other documents were presented in response to relevant queries with one or more of the query characteristics. The method further comprises: using the query-dependent measure for the email to determine a presentation characteristic for presenting an email search result that matches the email; and providing the email search result for presentation with the presentation characteristic in response to the query.
[0014] This method and other implementations of the technology disclosed herein may each optionally include one or more of the following features.
[0015] In some implementations, the at least one email attribute is based on both the sender content in the sender field and the subject content in the subject field. In some of these implementations, the at least one email attribute is a co-occurrence of the subject content in both the sender field and the subject content in the subject field. The sender content can contain a domain name of a sender's email address, and / or the subject content can contain a template that includes one or more terms and one or more wildcards.
[0016] In some implementations, at least one email attribute is based on the subject content in the subject field, and the subject content contains a template that includes one or more terms and one or more placeholders.
[0017] In some implementations, the other documents on which the measures are based exclude email.
[0018] In some implementations, the procedure further includes: generating a query-independent measure for the email based on additional measures of additional past interactions with the document attributes in response to additional queries that do not exhibit any of the query attributes; and furthermore, using the query-independent measures for the email to determine the presentation characteristic for presenting the email search result that corresponds to the email.
[0019] In some implementations, using the request-dependent measure for the email to determine the presentation characteristic involves the following: determining a score for the email based on the request-dependent measure; determining additional scores for other response documents; ranking the email relative to the other response documents based on the score and the additional scores; and determining the presentation characteristic based on the ranking.
[0020] In some implementations, the document attributes also include an email category. In some of these implementations, the procedure further involves using a machine learning model to determine the email category.
[0021] In some implementations, past interactions with other documents that exhibit one or more of the document characteristics include selections of the other documents.
[0022] In some implementations, a procedure is provided that involves receiving a request entered by a user via a user interface input device on the user's computing device and identifying response documents in response to the request. The response documents include access-restricted documents belonging to the user. These access-restricted documents can only be accessed by the user and any limited group of other users specified by the user.The procedure further comprises identifying one or more query characteristics for the query, and for each of several of the access-restricted documents: identifying one or more document characteristics for the access-restricted document; and generating a query-dependent measure for the access-restricted document based on measures of past interactions between the query characteristics and the document characteristics, wherein each of the measures is based on a set of past interactions by appropriate users with other documents exhibiting one or more of the document characteristics, when the other documents were presented in response to appropriate queries exhibiting one or more of the query characteristics, and wherein the other documents may optionally include several inaccessible documents to which the user cannot access.The procedure further includes using the request-dependent measures for the access-restricted documents to determine a presentation order for the response documents, and providing one or more of the response documents for presentation based on the presentation order in response to the request.
[0023] This method and other implementations of the technology disclosed herein may each optionally include one or more of the following features.
[0024] In some implementations, the document characteristics for the access-restricted document include a template that is contained in a specific field of the access-restricted document.
[0025] In some implementations, the other documents exclude one or more of the access-restricted documents.
[0026] In some implementations, the other documents on which a particular measure of the measures is based consist of inaccessible documents that the user cannot access.
[0027] In some implementations, the procedure further includes: for each of the access-restricted documents, generating a query-independent measure for the access-restricted document based on additional measures of additional past interactions with the document attributes in response to additional requests that do not exhibit any of the request attributes; and furthermore, using the query-independent measures for the access-restricted documents to determine the presentation order for the response documents.
[0028] In some implementations, a procedure is provided that includes receiving a request entered by a user via a user interface input device of a user computing device, identifying response documents in response to the request, and identifying one or more request characteristics for the request.The procedure further comprises, for each of several of the documents: identifying one or more document attributes for the document and generating a query-dependent measure for the document based on measures of past interactions between the query attributes and the document attributes, wherein each of the measures is based on a set of past interactions by corresponding users with other documents exhibiting one or more of the document attributes, when the other documents were presented in response to corresponding queries exhibiting one or more of the query attributes, and wherein the other documents contain multiple documents that exist in addition to the document.The procedure further includes using the request-dependent measures for the documents to determine a presentation order for the response documents, and determining one or more of the response documents to be presented based on the presentation order as a response to the request.
[0029] In some implementations, a procedure is provided that includes: selecting multiple document attributes and selecting multiple query attributes. Selecting each document attribute involves choosing the attribute based on its occurrence in access-restricted documents from at least a threshold set of users. Selecting each query attribute involves choosing the query attribute based on its occurrence in access-restricted queries from at least a threshold set of users. The access-restricted queries are those for which at least one of the access-restricted documents was provided as a response.The procedure further comprises, for each set of multiple tuples of query attribute and document attribute, each containing at least one of the query attributes and at least one of the document attributes: generating a past interaction measure between the query attributes and the document attributes of the query attribute, document attribute tuple. The generation of the past interaction measure is based on a set of past interactions with corresponding documents of the access-restricted documents, where the corresponding documents were presented in response to corresponding queries of the access-restricted queries, where the corresponding documents exhibit the document attributes of the query attribute, document attribute tuple, and where the corresponding queries exhibit the query attribute of the query attribute, document attribute tuple.The procedure further includes storing each of the past interaction measures in conjunction with a corresponding tuple query feature, document feature in one or more computer-readable media.
[0030] This method and other implementations of the technology disclosed herein may each optionally include one or more of the following features.
[0031] In some implementations, the procedure further includes: identifying a new document in response to a new request from a specific user, which comprises a new request group of document attributes; and generating a measure for the new document based on a group of past interaction measures. The group of past interaction measures can be selected based on the past interaction measures of the group stored in conjunction with the tuples `request attribute` and `document attribute`, each containing at least one of the document attributes of the new request group. The procedure further includes providing the new document in response to the new request based on the measure.In some of these implementations, the group of past interaction measures is further selected based on the past interaction measures of the group that are stored in conjunction with the tuples query attribute and document attribute, each containing at least one query attribute of the new query. In some implementations, the new document is excluded from the access-restricted documents that were used when generating the past interaction measures.
[0032] In some implementations, a procedure is provided that includes selecting multiple document attributes and selecting multiple query attributes. The procedure further includes, for each set of multiple query attribute and document attribute tuples, each containing at least one of the query attributes and at least one of the document attributes: generating a past interaction measure between the query attributes and the document attributes of the query attribute, document attribute tuple, where: generating the past interaction measure is based on a set of past interactions with corresponding documents when the corresponding documents were presented in response to corresponding queries; the corresponding documents have the document attributes of the query attribute, document attribute tuple; and the corresponding queries have the query attribute of the query attribute, document attribute tuple.The procedure further includes storing each of the past interaction measures in conjunction with a corresponding tuple query feature, document feature in one or more computer-readable media.
[0033] Other implementations may include one or more non-transient, computer-readable storage media that hold instructions executable by one or more processors to perform a procedure such as one or more of the procedures described herein. Yet another implementation may include a system comprising memory and one or more processors capable of executing instructions stored in the memory to perform a procedure such as one or more of the procedures described herein.
[0034] It should be noted that all combinations of the preceding concepts and additional concepts described in more detail herein are considered part of the subject matter disclosed herein. For example, all combinations of claimed subjects appearing at the end of this disclosure are considered part of the subject matter disclosed herein. Brief description of the drawings Fig. Figure 1 is a block diagram of an exemplary environment in which some of the implementations revealed here can be implemented. Fig. Figure 2A shows a representation of part of a request document model according to different implementations. Fig. Figure 2B shows a representation of part of a query feature model according to different implementations. Fig. Figure 2C shows a representation of part of a document feature model according to different implementations. Fig. Figure 2D shows a representation of part of a query feature-document feature model according to different implementations. Fig. Figure 3 shows an example of multiple access-restricted documents from multiple users. Fig. Figure 4 is a flowchart that shows an exemplary procedure for generating past interaction measures for each of several tuples query feature, document feature. Fig. Figure 5 is a flowchart illustrating an exemplary procedure for using query-dependent and / or query-independent measures of documents in response to a query to classify the documents and to provide the appropriate search results based on the classification. Fig. Figure 6 illustrates an exemplary client computing device with a display screen that shows access-restricted search results with presentation characteristics determined according to implementations disclosed herein. Fig. Figure 7 shows an exemplary architecture of a computing device. Detailed description
[0035] Some implementations disclosed herein may be applicable to access to access-restricted documents. As used herein, an “access-restricted document” is contrasted with a publicly accessible document (e.g., freely available to the public via the World Wide Web) and is an electronic document that can be accessed by a restricted group of users. In some implementations, access to an access-restricted document may be limited to the restricted group of users based on their login credentials, based on the fact that the access-restricted document is accessible via a private network to which only the restricted group of users has access, and / or based on other techniques.As used herein, a "user-restricted document" is a document with restricted access that can be accessed only by the user and, optionally, a limited group of one or more other users, as specified by the user or otherwise controlled. For example, a user-restricted document may be accessible only to the user depending on: whether it is stored locally on a computing device controlled by the user; whether it can be accessed through one or more computer applications using appropriate user credentials; and so on. For example, a user's emails may be user-restricted documents that can be accessed only by the user using appropriate user credentials.For example, a user's heterogeneous documents stored in a cloud-based storage system can be access-restricted documents, accessible only to that user with appropriate login credentials. Optionally, one or more of these heterogeneous documents can also be accessed by a limited group of other users, based on explicit authorization from the user via one or more computer applications. Similarly, various documents stored locally on a user's mobile phone, tablet, desktop, and / or other computing devices can be access-restricted documents because they are stored locally on the user's one or more computing devices.
[0036] User interaction data (e.g., a click-through rate) has been used to rank specific publicly available search result documents for particular queries. For example, user interaction data might indicate that for a specific search query, a particular publicly available search result document responding to that query has a click-through rate that is significantly higher than that of other publicly available search result documents responding to that query. Based on such an indication, a search result corresponding to that specific publicly available search result document might be ranked more prominently for that query (e.g., given priority for display) than search results for the other publicly available responding search result documents.
[0037] However, different techniques that rely on using user interaction data to rank publicly available search results for specific queries may not be applicable to different documents and / or may not provide the desired performance. For example, different techniques may not work for different access-restricted documents (e.g., documents accessed by a user who submits a query) and / or different publicly available documents (e.g., publicly available documents that have no and / or relatively few interactions in response to queries).
[0038] As an example, suppose a user submits a search query to search their personal email, and several response emails (i.e., access-restricted documents belonging to the user) are identified as answers to the query (e.g., the emails contain one or more terms that match one or more terms in the search query). It may be the case that one or more (e.g., all) of the response emails have never been presented as answers to previous searches by other users and / or the user, and / or the user has never interacted with them. For example, a particular email might be one that was sent only to the user and with which the user has never previously interacted in response to a search query.Accordingly, there may be no user interaction data associated with the specific email, rendering various techniques relating to the use of user interaction data to rank publicly available search results ineffective for ranking that specific email.
[0039] As another example, suppose a user performs a search query to search a corpus of access-restricted documents available to a limited group of users, and that several response documents are identified as answers to the query. It may be the case that one or more (e.g., all) of the response documents were never presented in response to previous submissions of the query and / or were never interacted with, and / or were presented in response to previous submissions of the query only in negligible quantities and / or were interacted with only in negligible quantities.Accordingly, there may not be enough user interaction data in response to the search query associated with such documents, rendering various techniques relating to the use of user interaction data to rank publicly available search results ineffective for ranking such documents.
[0040] As a further example, suppose a user submits a search query to search a corpus of publicly available documents, and several response documents are identified as answers to the query. It may be the case that one or more (e.g., all) of the response documents were never presented in response to previous submissions of the search query, and / or were never interacted with, and / or were presented in response to previous submissions of the search query only in negligible quantities, and / or were interacted with only in negligible quantities.Accordingly, there may not be enough user interaction data in response to the search query associated with such documents, rendering various techniques relating to the use of user interaction data to rank publicly available search results ineffective for ranking such documents.
[0041] This description introduces various technical features related to using one or more document features of a document in response to a query, and optionally one or more query features of the query, to determine a presentation characteristic for presenting a search result that corresponds to the document—and, in response to the query, providing the search result for presentation with the presentation characteristic. Measures associated with the one or more document features and / or the one or more query features can be used to determine the presentation characteristic. The measures can be based on past interactions of relevant users with other documents that share one or more of the document features with the document, where several of the other documents are each distinct from the document (and optionally distinct from each other).Using such measures makes it possible to leverage past interactions with other documents when determining interaction-based relevance of the access-restricted document, optionally without reference to any past interactions specifically directed at the access-restricted document. In some implementations, the other documents contain documents that are themselves access-restricted or are restricted to them.
[0042] In some implementations, when determining the presentation characteristic of a search result corresponding to a document in response to a query, a query-dependent measure of the access-restricted document is generated and used to determine the presentation characteristic. In some of these implementations, the query-dependent measure is based on past interaction measures between query features of the query and document features of the document. Each of the measures can be based on a set of past interactions by corresponding users with other documents that exhibit one or more of the document features, when the other documents were presented in response to corresponding queries with one or more of the query features.
[0043] As an example, suppose a user uses an email search interface to submit a request for a "book order number." A corpus of the user's emails, each containing access-restricted documents, can be searched, and multiple reply emails can be identified as responding to the request. A specific reply email might originate from "store@exampleurl.com," have "Order Confirmation 1A2B3C" as the subject, and contain body text identifying a specific book purchased by the user, along with purchase details (e.g., purchase date, shipping address, delivery date, cost). This specific reply email may never have been interacted with by other users in response to requests from other users (i.e.,(since it is personal to the user and inaccessible to other users) – and the user may never have interacted with it in response to a query. However, the techniques described herein can still be used to determine a query-dependent measure for the specific email based on past interaction measures between query features of the "book order number" query and document features of the specific email.
[0044] For example, a first past interaction measure can be determined based on a set of interactions from multiple users with other emails that have "store@exampleurl.com" in a sender field and "Order Confirmation [#]" (where [#] is a wildcard indicating an alphabetic and / or numeric string) in a subject field, if these other emails were presented in response to corresponding requests containing N-grams of "book order". Similarly, a second interaction measure can be determined based on a set of interactions from multiple users with other emails that contain "store@exampleurl.com" in a sender field and "Order Confirmation [#]" in a subject field, if these other emails were presented in response to corresponding requests containing N-grams of "order number".The query-dependent measure can be generated based on the first measure, the second measure, and optionally other similarly determined measures. For example, the query-dependent measure can be a sum, an average, a median, or another statistical combination of the measures.
[0045] The query-dependent measure can be used to determine a presentation characteristic for the specific response email. For example, the query-dependent measure can be used to modify an initial score for the specific response email (e.g., a score based on a degree of match between the query and the specific email), and the score can be used to rank the specific email relative to other response emails (e.g., based on optionally modified initial scores for those emails).The classification can be used, for example, to determine which reply emails are initially used to present relevant search results in response to the query, to determine a presentation order (or other display prominence) for the search results, and / or to determine additional or alternative presentation characteristics for search results.
[0046] In some implementations, a query-independent measure for the document is generated and used additionally or alternatively to determine its presentation characteristics. In some of these implementations, the query-independent measure is based on measures of past interactions by relevant users with other documents that exhibit one or more of the document's features, when the other documents were presented in response to relevant queries. These queries include or are limited to queries that do not contain any of the query features. Accordingly, the query-independent measure can provide an indication of the general popularity of documents with the one or more document features, while the query-dependent measure provides an indication of the popularity of documents with the one or more document features in response to queries containing the query features.
[0047] In some implementations, a query-dependent measure and / or a query-independent measure of a document can be generated based on a query feature-document feature model. The query feature-document feature model can be generated based on a query-document model, a document feature model, and / or a query feature model.
[0048] The query-document model can, for example, be a bipartite graph that models the interactions between queries and documents, as indicated by one or more stored records of past queries and corresponding interactions. For instance, the nodes of the query-document graph can represent queries and documents. The edges can lie between query and document nodes and can, for example, represent whether the corresponding document was observed for the corresponding query (e.g., a relevant search result was presented as a response to the query) and / or whether the document was interacted with for the corresponding query (e.g., a relevant search result was selected).
[0049] The document feature model can, for example, be a bipartite graph that models the relationship between documents and their document features. Various features, such as category features, structural features, and / or n-gram features, can be used. For example, a document's category features can specify one or more categories to which the document belongs. This can be based, for instance, on applying document features to a classifier or other machine learning model and determining the category features based on the output generated by the machine learning model. As an example of categories, emails might belong to finance, travel, order confirmations, and / or other categories. Structural features can specify templates and / or other content of certain structural fields within documents.For example, structural features for emails or other electronic communications can include: sender content contained in a sender field of the electronic communication (e.g., a domain name of the sender's email address, a relationship between the sender and the user), subject content contained in a subject field of the electronic communication (e.g., a specific template to which the subject field corresponds, such as "Order Confirmation [#]"), and / or a co-occurrence of specific sender content and specific subject content (i.e., the sender content and the subject content both appear in their respective fields). For example, structural features of an access-restricted document can include a file type feature based, for instance, on a file extension of the access-restricted document. Other structural features can include content such as...Templates and / or N-grams that appear in one or more specific additional and / or alternative fields of a document, such as a document title field, a title, location and / or notes field of a calendar entry document, etc.
[0050] The query feature model can, for example, be a bipartite graph that models the relationship between queries and their query features. The query features of a query can include, for example, one or more n-grams appearing in the query (e.g., the longest n-gram appearing in the query), one or more entities referenced in the query (e.g., a specific person, place, and / or object), one or more entity types referenced in the query (e.g., a city, a person's name, a location, a restaurant), grammatical features of the query, and so on.
[0051] The query feature-document feature model can, for example, be a bipartite graph generated using the query-document graph, the document feature graph, and the query feature graph. The query feature-document feature model models the interactions between document features and query features. In other words, it models interactions between document and query features, rather than interactions directly between queries and documents. In some implementations, it is generated based on a transformation of the query-document model into the "document feature" and "query feature" spaces, which are jointly modeled by the document feature model and the query feature model.
[0052] In many implementations, only features (query features or document features) that occur with at least a threshold frequency (in queries or documents) and / or for at least a threshold number of users can be used when generating the query-feature, document-feature, and / or query-feature-document-feature graphs. In some of these implementations, this can ensure that features do not contain sensitive information by guaranteeing that these features occur with at least a threshold frequency and / or for at least a threshold number of users.
[0053] The query feature-document feature model can be used to determine a query-independent measure and / or a query-dependent measure for a given document. For example, to determine a query-dependent measure for a given query that has one or more given query features, edges between the given query features and document features of the given document can be determined. Each edge provides a measure of past interactions between a corresponding query feature and a corresponding document feature. The measures can be combined (e.g., summed and / or another statistical combination) to determine the query-dependent measure. For example, to determine a query-independent measure for the given document, edges between all query features and document features of the document can also be determined. The measures can be combined (e.g., summed and / or another statistical combination) to determine the query-dependent measure.summed and / or another statistical combination) to determine the query-independent measure.
[0054] With reference to Fig. Figure 1 presents an exemplary environment in which the techniques disclosed herein can be implemented. The exemplary environment includes a client device 106, a search system 110, a past interaction measure system 120, and a document measure system 130. The exemplary environment further includes one or more personal corpora 158 of a user of the client device 106. The one or more personal corpora 158 may each be stored on one or more corresponding non-transient, machine-readable media, which may be located on the client device 106 and / or remotely from the client device 106 (e.g., on one or more remote servers). The one or more personal corpora 158 may each store one or more access-restricted documents of the user, such as the user's electronic communications (e.g., emails, letters, etc.).Emails, SMS messages, chat messages, messages from social networks), media files (e.g. audio files, image files, video files), word processing documents, calendar entries, contact entries, etc.
[0055] The exemplary environment further includes a query-document model 150, which may be stored on one or more non-transient, machine-readable media. The query-document model 150 may, for example, be a bipartite graph that models the interactions between queries and documents (including or limited to access-restricted documents) as specified by one or more stored records of past queries and corresponding interactions. For example, the query-document model 150 may be generated based on records of past queries and corresponding interactions provided by the search system 110 and / or other search systems based on interactions with the search system by multiple users via multiple corresponding client devices.The exemplary environment further includes one or more additional models 160, which are generated by the past interaction measure system 120 and can be used by the document measure system 130. For example, the one or more additional models 160 can comprise at least one query feature-document feature model.
[0056] A user of the client device 106 can submit queries to the search system 110 via one or more user interface input devices of the client device 106. For example, the user can speak the query using a microphone of the client device 106, type the query using a hardware keyboard and / or virtual keyboard of the client device 106, etc. In response to a query from the client device 106, the search system 110 searches the user's one or more personal corpora 158 to identify access-restricted documents of the user in response to the search query (if any), using, for example, conventional and / or other information retrieval techniques.In some implementations, the one or more personal corpora 158 may contain an index that indexes documents within it based on one or more attributes, and the search system 110 identifies response documents using the index. In some implementations, the search system 110 additionally or alternatively searches one or more corpora that contain or are restricted to one or more access-restricted documents that are not user-restricted documents and / or publicly accessible documents.
[0057] The search system 110 includes a rating engine 112 that calculates scores for the documents identified in response to a search query, for example, using one or more rating signals. Each rating signal provides information about the document itself and / or the relationship between the document and the search query.
[0058] In many implementations, the rating signals on which the Rating Engine 112 calculates scores for a given document include a query-dependent measure and / or a query-independent measure generated by the Document Measure System 130 according to the implementations described here. In some implementations, the Rating Engine 112 may use additional rating signals, such as rating signals that indicate a degree of match between the given document and the search query. For example, the rating signals for a document may be based on whether each of one or more query terms appears in the document, where each of one or more query terms appears in the document, the term frequency of each of one or more of the query terms that appear in the document, and / or the document frequency of each of one or more of the query terms that appear in the document.
[0059] The 112 Ranking Engine then ranks the response documents using the scores. The 110 Search System uses the response documents ranked by the 112 Ranking Engine to generate search results provided in response to the query. The search results contain results that correspond to the documents in response to the search query. For example, each of one or more search results may contain a title of a respective document, a link to a respective document, and / or a summary of the content of a respective document. For example, the content summary may include a specific snippet or section of the document in response to the search query.Furthermore, for example, a search result associated with an image document may include a thumbnail of the image, a title associated with the image, and / or a link to the image. Similarly, for example, a search result associated with a video document may include an image from the video, a segment of the video, the video title, and / or a link to the video. Other search results include a summary of information in response to the search query. This summary may be generated from one or more documents related to the search query and / or from other sources.
[0060] The search results are provided in a format that allows them to be presented to the user via one or more user interface output devices of the client device 106 (e.g., a display and / or a speaker). For example, the search results can be presented by the client device 106 in one or more pop-up windows or other interfaces rendered in an application running on the client device 106, and / or as one or more search results delivered to a user via audio. Fig. Figure 6 presents an example of client device 106, which displays search results, and is described in further detail herein. The search results can be presented to the user with one or more presentation characteristics based on the ranking of the corresponding search result documents. For example, the most prominently displayed search result can be the highest-ranked search result, the second most prominently displayed search result can be the second-highest-ranked search result, and so on. For example, only a subset of all search results can be initially presented, and this subset can be selected based on the ranking.
[0061] The client device 106 can be, for example, a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device in the user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), or a user-worn device containing a computing device (e.g., the user's watch with a computing device, the user's glasses with a computing device). Additional and / or alternative client devices may be provided. The client device 106 typically includes one or more applications to facilitate the submission of search queries and the sending and receiving of data over a network.
[0062] Although they are in Fig. Although the search system 110 is shown separately in Figure 1, in some implementations it can be implemented wholly or partially by the client device 106. For example, the one or more personal corpora 158 may contain documents stored locally on the client device 106, and the search system 110 can search such locally stored documents. In some implementations, the search system 110 can be implemented wholly or partially by one or more remote computing devices, and the client device 106 can communicate with the search system 110 over a network such as a local area network (LAN) or a wide area network (WAN) (e.g., the Internet). Fig. 1. While the search system 110 is configured to be connected only to the client device 106 and only to the one or more personal corpora 158 containing access-restricted documents of a user of the client device 106, in some implementations the search system 110 can be connected to multiple client devices and / or access multiple corpora, such as one or more personal corpora of multiple users and / or non-personal corpora. For example, the search system 110 can be an email search system of an email service and can search a first user's personal email corpus in response to requests from the first user, search a second user's personal email corpus in response to requests from the second user, and so on.
[0063] Although only a single search system 110 in Fig. As shown in Figure 1, multiple search systems 110 can be provided, and each can use query-dependent and / or query-independent measures provided by the document measure system 130 (or a separate instance thereof). Even if the search system 110 is shown to communicate only with the one or more personal corpora 158, in some implementations the search system 110 can additionally or alternatively search one or more non-personal corpora, such as one or more public corpora and / or one or more non-personal corpora that search access-restricted documents. For example, the search system 110 can additionally search one or more public corpora and return search results that include both public and access-restricted content. Even if the document measure system 130 in Fig. Although document measurement system 130 is represented as separate from search system 110, some implementations allow one or more aspects of both to be combined in a single system. For example, in some implementations, one or more aspects of document measurement system 130 can be implemented by the classification engine 112 of search system 110.
[0064] In some implementations, the document measure system 130 may include a document feature engine 132, a query feature engine 134, a query-dependent measure engine 136, and / or a query-independent measure engine 138. In some implementations, all engines 132, 134, 136, and / or 138, or aspects thereof, may be omitted, combined, and / or implemented in a component separate from the document measure system 130.
[0065] The document measure system 130 receives from the search system 110 a reference to a query that has been transmitted to the search system 110, and / or a reference to one or more documents that have been identified by the search system 110 in response to the query, such as access-restricted documents from the one or more personal corpora 158.
[0066] The document attribute engine 132 identifies one or more document attributes for each document. Various document attributes, such as category attributes, structural attributes, and / or n-gram attributes described herein, can be identified. For example, for an image document, in response to a query, document attributes may include n-grams or other information specifying one or more particular objects and / or classes of objects present in the image document (e.g., based on automated image analysis and / or human-applied tags). In some implementations, the entire document attribute engine 132, or aspects thereof, may be implemented by the search system 110.
[0067] The query feature engine 134 identifies one or more query features for the query. Various query features can be identified, such as one or more n-grams appearing in the query, one or more entities referenced in the query, one or more entity types referenced in the query, grammatical features, etc. In some implementations, the entire query feature engine 134, or aspects thereof, may be implemented by the search system 110.
[0068] The Query-Dependent Measures Engine 136 generates a query-dependent measure for each of the documents. When determining a query-dependent measure for a document, the Query-Dependent Measures Engine 136 determines past interaction measures that are assigned in Model 160 to the query features and document features determined by Engines 132 and 134. For example, suppose query features QF1 and QF2 for a query (where QF denotes a query feature) and document features DF1, DF2, and DF3 for an access-restricted document in response to the query (where DF denotes a document feature). The Query-Dependent Measures Engine 136 can determine a past interaction measure for each of the following: QF1 and DF1, QF1 and DF2, QF1 and DF3, QF2 and DF1, QF2 and DF2, and QF2 and DF3.The query-dependent measure engine 136 can then generate the query-dependent measure for the access-restricted document based on a combination of the six separate past interaction measures.
[0069] Each of the past interaction measures used by the query-dependent measure engine 136 can be based on a set of past interactions by relevant users with other documents that exhibit one or more of the document characteristics, provided that the other documents were presented in response to relevant queries with one or more of the query characteristics. The other documents themselves may contain or be restricted to multiple access-restricted documents, such as inaccessible documents, each of which is personal to and inaccessible to a relevant user. An additional description of the generation of past interaction measures is provided herein.
[0070] The query-independent measure engine 138 generates a query-independent measure for each document. When determining a query-independent measure for a document, the query-independent measure engine 138 determines past interaction measures that are assigned in model 160 to a set of query features and the document features determined by the engine 134. The set of query features contains query features that are in addition to, or limited to, those determined by the query feature engine 134. Accordingly, the set of query features is independent of the query to which the document is a response, in the sense that it contains query features in addition to the query features of the query. As an example, consider the document features DF1, DF2, and DF3 for an access-restricted document (where DF denotes a document feature).The query-independent measure engine 138 can determine: all past interaction measures between the group of query features and DF1, all past interaction measures between the group of query features and DF2, and all past interaction measures between the group of query features and DF3. Suppose the query feature group contains query features QF1-QF1000. For DF1, past interaction measures can be determined for QF1 and DF1, QF2 and DF1, QF3 and DF1, ... and QF1000 and DF1. The query-independent measure engine 136 can then generate the query-dependent measure based on a combination of the past interaction measures.
[0071] The document measure system 130 provides the query-dependent measure and / or the query-independent measure for each document to the search system 110. The ranking engine 112 can use the query-dependent measures and / or the query-independent measures when ranking the documents and can use the ranking to determine a presentation order and / or other presentation characteristics for search results for the documents. In some implementations, the ranking engine 112 uses the query-dependent measure and / or the query-independent measure to determine a score for the document and uses the score to rank the document. For example, the ranking engine 112 can adjust a base score for the document (e.g., a base score based on other ranking signals) in light of the query-dependent measure and / or the query-independent measure to produce a modified score.
[0072] As an example, let's consider a base score of sc b for a document accepted for a request. This base score can be based, for example, on a keyword comparison and / or one or more other rating signals. The rating engine 112 can assign a final score sc f based on f(sc b , M d , M q,d ) determine where M d represents the query-dependent measure for the document, and where M q,d The query-independent measure for the document is represented. f(·) can optionally be a manually agreed-upon score or a machine-learned rating function. In some implementations, the base score 112 holds the base score (sc b ) fixed and trains an adaptation δ(M d , M q,d ) via the base score (sc b The valuation function f(·) thus becomes: f(sc b , M d , M q,d ) = sc b + δ(Md , M q,d ). This adaptive formulation can be advantageous for environments where the base score is already highly optimized and optionally separate from the query-independent and / or query-dependent measure.
[0073] In some implementations, the Past Interaction Measure System 120 may include a Query Document Model Engine 122, a Document Feature Model Engine 124, a Query Feature Model Engine 126, and / or a Query Feature Document Feature Model Engine 128. In some implementations, all of the engines 122, 124, 126, and / or 128, or aspects thereof, may be omitted, combined, and / or implemented in a component separate from the Past Interaction Measure System 120.
[0074] The query-document model engine 122 generates the query-document model 150. In some implementations, the entire query-document model engine 122, or aspects of it, may be implemented by the search system 110. The query-document model 150 may, for example, be a bipartite graph that models the interactions between queries and documents, as indicated by one or more stored records of past queries and corresponding interactions. For example, the nodes of the query-document graph may represent queries and documents. The edges may lie between query and document nodes and may, for example, represent whether the corresponding document was observed for the corresponding query (e.g., a corresponding search result was presented as a response to the corresponding query) and / or whether the document was interacted with for the corresponding query (e.g.,(Selection of a corresponding search result). In some implementations, each edge can contain a binary representation of whether an interaction occurred. In some implementations, the edges can be weighted based on a type of interaction. For example, selecting a search result followed by accessing the underlying document for at least a threshold duration can be weighted more heavily than selecting the underlying document followed by accessing it for less than the threshold duration, and this in turn can be weighted more heavily than "circling" the search result without making a selection.
[0075] In some implementations, the request-document model 150 can be replaced by a triple (Q,D,EQD) be represented, whereby Q the set of query nodes is that which represent corresponding queries, D is the set of document nodes that represent corresponding documents, and the edge set EQD The edges represent the connections between the query nodes and document nodes. The edges in the edge set EQD can be defined by tuples of the form e(q, d) = < γ o (q, d), γ c (q, d) > can be parameterized, where q represents a query node connected by the edge, d represents a document node connected by the edge, and parameterization functions γ o (a, b) and γ c (a, b) indicate that entities a and b were observed or clicked in the same search session.
[0076] In this description, the term "graph" is used generally to refer to any representation of multiple associated information elements. A graph, or a portion of a graph, need not reside in a single storage device and may contain pointers or other references to information elements that may reside in other storage devices. For example, a graph may contain multiple nodes that are associated with one another, each node containing an identifier of an entity or other information element that may reside in a different data structure and / or storage medium.
[0077] The document feature model engine 124 generates a document feature model, which can optionally be included in one or more models 160. The document feature model engine 124 can generate document features based on documents contained in the request document model. For example, the engine 124 can identify one or more document features for each of the documents in the request document model 150 and define a relationship between the document and its document features. The document feature model can, for example, be a bipartite graph that models the relationship between documents and their document features. For example, a first node in the model can represent a document feature, and this node can be connected by corresponding edges to each of several document nodes, each representing a corresponding document containing the document feature.The edges can each indicate whether a corresponding feature is present in a corresponding document and, optionally, specify a weight for that feature within the document (e.g., the weight for a category feature can indicate how strongly the document is assigned to that category). Various features, such as category features, structural features, and / or N-gram features, can be used.
[0078] In some implementations, the document feature model can be represented by a triple. (D,AD,ED) be represented, whereby D the set of document nodes that represent corresponding documents, where A D the set of document feature nodes that represent the set of document features, and the edge set ED The edges that connect the document nodes and the document feature nodes are represented. The edges in the edge set ED can through e(d,aijd) be parameterized, whereby e(d,aijd) indicates whether a corresponding feature is present in a corresponding document, and, if applicable, indicates a weighting of the corresponding feature for the corresponding document.
[0079] The query feature model engine 126 generates a query feature model, which can optionally be included in one or more models 160. The query feature model engine 126 can generate the features for queries that are contained in the query document model 150. For example, the engine 126 can identify one or more query features for each of the queries in the query document model 150 and define a relationship between the query and its query features. The query feature model can, for example, be a bipartite graph that models the relationship between queries and their query features. For example, a first node in the model can represent a query feature, and this node can be connected by corresponding edges to each of several query nodes, each representing a corresponding query that contains the query feature.The edges can each indicate whether a corresponding feature is present in a corresponding query and, if applicable, specify a weighting of the corresponding feature for that query. Various features can be used, such as one or more n-grams appearing in the query, one or more entities referenced in the query, one or more entity types referenced in the query, grammatical features of the query, etc.
[0080] The query feature model can be described by a triple (Q,AQ,EQ) be represented, whereby Q the set of query nodes that represent corresponding queries, where AQ the set of query feature nodes that represent the set of query features, and the edge set EQ The edges that connect the query nodes and the query feature nodes are represented. The edges in the edge set EQ can through e(q,aklq) be parameterized, whereby e(q,aklq) indicates whether a corresponding query characteristic is present in a corresponding query, and, if applicable, specifies a weighting of the corresponding characteristic for the corresponding query.
[0081] The query feature-document feature model engine generates a query feature-document feature model, which can optionally be contained within one or more models. The query feature-document feature model can, for example, be a bipartite graph generated using the query-document graph, the document feature graph, and the query feature graph. The query feature-document feature model models the interactions between document features and query features. In other words, it models interactions between document and query features, rather than interactions directly between queries and documents. In some implementations, it is generated based on a transformation of the query-document model into the "document feature" and "query feature" space, which is jointly modeled by the document feature model and the query feature model.
[0082] The query feature-document feature model can be described by a triple (AQAD,EA) be represented, whereby AQ the set of query feature nodes that represent the set of query features, A D the set of document feature nodes that represent the set of document features, and the edge set ℇ A The edges represent those that connect the query feature nodes and the document feature nodes. The edges in the edge set ℇ A Each edge has a weight or other measure based on the number of past interactions between the query feature of the corresponding query feature node and the document feature of the corresponding document feature nodes. The edges in the edge set E ^ A can be parameterized as follows: e(aklq,aijd)=∑q:e(q,aklq)=1∑d:e(d,aijd)=1e(q,d) =<γo(aklq,aijd),γc(aklq,aijd)>, where the edge functions e(·) are defined as described above. As can be understood by considering the parameterization of the edges outlined above, the parameterization models a query-document attribute of observed and clicked associations by summing all queries and documents that can be assigned to their respective attributes.
[0083] In many implementations, only features (query features or document features) that occur with at least a threshold frequency (in queries or documents) and / or for at least a threshold number of users can be used when generating query-feature, document-feature, and / or query-feature-document-feature models. In some of these implementations, this can ensure that feature nodes do not contain sensitive information by guaranteeing that features of these feature nodes occur with at least a threshold frequency and / or for at least a threshold number of users.In some of these implementations, this can be achieved by removing all document feature nodes from the document feature graph that do not have at least a threshold number of edges indicating their presence in corresponding documents; and / or by removing all query feature nodes from the query feature model that do not have at least a threshold number of edges indicating their presence in corresponding queries. Additionally or alternatively, query feature nodes and / or document feature nodes can be removed from the query feature-document feature model using similar techniques.
[0084] The query feature-document feature model can be used to determine a query-independent measure and / or a query-dependent measure for a given document. For example, to determine a query-dependent measure for a given query that has one or more given query features, edges can be determined between the one or more given query features and document features of the document. Each edge provides a measure of past interactions between a corresponding query feature and a corresponding document feature. The measures can be combined (e.g., summed and / or another statistical combination) to determine the query-dependent measure.Furthermore, to determine a query-independent measure for the given document, edges between a group of query features (which includes or is limited to query features not contained in the given query features) and document features of the document can be determined. The measures can be combined (e.g., summed and / or another statistical combination) to determine the query-independent measure.
[0085] An additional description of different models that can be used in different implementations is provided with reference to Fig. 2A - 2D and 3 provided.
[0086] Fig. Figure 2A shows a representation of a section of the request-document model 158 according to different implementations. The section contains a request node 152A, which is connected to a document node 153A by an edge 151A. The section also contains a separate request node 152B, which is connected to another document node 153B by an edge 151B.
[0087] The query node 152A represents a specific query, and the document node 153A represents a specific document. For the purposes of a working example, it is assumed that query node 152A represents a query for "book order number" and document node 153A represents email 353A from Fig. 3 represents. Edge 151A represents that a user interacted with a search result for the document corresponding to document node 153A in response to the query corresponding to query node 152A.
[0088] Query node 152B represents a specific query that differs from the specific query represented by query node 152A, and document node 153A represents a specific document that differs from the one represented by document node 153A. For the purposes of this working example, query node 152B is assumed to represent a "book order" query, and document node 153B is assumed to represent email 353B from Fig. 3 represents. Edge 151B represents that a user interacted with a search result for the document corresponding to document node 153B in response to the query corresponding to query node 152B.
[0089] It is understood that the request-document model 158 will contain a large number of additional request nodes, document nodes, and edges. For example, additional edges will be provided that connect additional request nodes and additional document nodes. It is also possible, for example, that additional edges may be connected to one or more of the nodes 152A, 152B, 153A, and 153B. For example, the document represented by document node 153A may have been selected in response to several different requests. Furthermore, for example, the request represented by request node 152A may have been issued by multiple users and used to select several different documents, such as multiple access-restricted documents belonging to those users.
[0090] Fig. Figure 2B shows a representation of a section of a query feature model 160A of one or more models 160 according to different implementations. Query node 152A is connected to query feature nodes 162A-C by corresponding edges 161A-C, indicating that the query represented by query node 152A has the query features represented by query feature nodes 162A-C. Query node 152B is connected to query feature nodes 162A and 162C by corresponding edges 161A and 161C, but not to query feature node 162B. The absence of an edge between query node 152B and query feature node 162B indicates that the query represented by query node 152B does not have the query feature represented by query feature node 162B. In some implementations, an edge may still be defined, but it may be indicated that the feature is not present (e.g.,This edge can be a “non-present” edge, whereas edges 161C and 161D can be “present” edges.
[0091] With further reference to the working example, the query feature node 162A can be a query feature of an N-gram "book order", the query feature node 162B can be a query feature of an N-gram "book order number", and the query feature node 162C can be a query feature of an N-gram "order".
[0092] It is understood that the query feature model will contain a large number of additional query nodes, query feature nodes, and edges. For example, additional query feature nodes can be connected to each of query nodes 152A and 152B. Furthermore, each of query feature nodes 162A-C can be connected to multiple additional query nodes. Additionally, these additional query nodes and query feature nodes will be provided with corresponding edges.
[0093] Fig. Figure 2C shows a representation of a section of a document feature model 160B of one or more models 160 according to different implementations. Document node 153A is connected to document feature nodes 164A and 164B by corresponding edges 163A and 163B, indicating that the document represented by document node 153A has the document features represented by document feature nodes 164A and 164B. Document node 153B is connected to document feature nodes 164A and 164C by corresponding edges 163C and 163D, indicating that the document represented by document node 153B has the document features represented by document feature nodes 164A and 164C. The absence of an edge between document node 153A and document feature node 164C indicates that the document represented by document node 153A does not have the document feature represented by document feature node 164C.Similarly, the absence of an edge between document node 153B and document feature node 164B indicates that the document represented by document node 153B does not have the document feature represented by document feature node 164B. In some implementations, edges may still be defined instead of being missing, but they may indicate that the corresponding document feature is not present.
[0094] With further reference to the working example, document attribute node 164A can be a structural document attribute, such as one that specifies certain content in a From field and / or a Sender field, as present in emails 353A and 353B. For example, document attribute node 164A can specify the simultaneous occurrence of the domain name "@exampleurl.com" in a Sender field and the template "Purchase Confirmation - [#]" in a Subject field, where [#] is a wildcard indicating an alphabetic and / or numeric string. As another example, document attribute node 164A (or an additional document attribute node) can instead specify a co-occurrence of certain content in both a Sender field and a Subject field (e.g., a co-occurrence of "store@exampleurl.com").Document feature node 164A can specify an N-gram from the body of email 353A, such as the fictitious book title "Soon Potter". Document feature node 164C can specify an N-gram from the body of email 353A, such as the fictitious book title "Fear and Dislike in Los Angeles".
[0095] It is understood that the document feature model will contain a large number of additional document nodes, document feature nodes, and edges. For example, additional document feature nodes can be connected to each of document nodes 153A and 153B. Similarly, each of document feature nodes 164A-C can be connected to multiple additional document nodes. Furthermore, additional document nodes and document feature nodes will, for example, be equipped with corresponding edges.
[0096] Fig. Figure 2D shows a representation of a section of a query feature-document feature model 160C of one or more models 160 according to different implementations. The query feature nodes 162A-C are each connected to corresponding document feature nodes 164A-C by corresponding edges. The edges of Fig. For simplicity, the 2D diagrams are not labeled. Each of the edges of Fig. 2D defines a corresponding past interaction measure between a corresponding query feature node and a document feature node and can be generated based on the models that are partly shown in 2A-2C.
[0097] It should be noted that when generating the past interaction measures, which are defined by the edges of Fig. 2D are defined, the two request-to-document interactions that are in Fig. The elements shown in Figure 2A will have a positive influence on the past interaction measures of all edges except the edge between query feature node 162B and document feature node 164C. This is because, as shown by the models of Fig. 2A - 2C indicate the interactions that occur in Fig. Figure 2A shows no interaction between query feature node 162B and document feature node 164C. In other words, since document feature node 164C is associated with document node 153B but not with document node 153A, and query feature node 162B is associated with query node 152A but not with query node 152B, the interactions of Fig. No interactions between query feature node 162B and document feature node 164C are assigned to 2A. As with the other models, it is understood that the query feature-document feature model will contain a large number of additional query feature nodes, document feature nodes, and edges.
[0098] Fig. Figure 4 is a flowchart illustrating an exemplary Procedure 400 for generating past interaction measures for each of several tuples of query feature and document feature. For simplicity, the operations of the flowchart are described with reference to a system that performs the operations. This system may include various components from different computer systems, such as one or more components of the Past Interaction Measure System 120. Although the operations of Procedure 400 are shown in a specific order, this is not intended to be restrictive. One or more operations may be rearranged, omitted, or added.
[0099] In block 452, the system selects several document attributes. For example, the system can select document attributes based on whether they occur in access-restricted documents used by at least a threshold number of users and / or in at least a threshold number of documents. In some implementations, the system selects document attributes based on those attributes that are features of documents contained in a query-document model, as described herein. In some implementations, when selecting document attributes, the system generates a document attribute model, as described herein.
[0100] In block 454, the system selects several query characteristics. For example, the system can select query characteristics based on the fact that these characteristics appear in queries for access-restricted documents from at least a threshold number of users and / or in at least a threshold set of such queries. In some implementations, the system selects query characteristics based on those characteristics that are features of queries contained in a query-document model, as described herein. In some implementations, when selecting query characteristics, the system generates a query-characteristic model, as described herein.
[0101] In block 456, the system selects a tuple query attribute, document attribute. For example, the tuple query attribute, document attribute can be a single query attribute and a single document attribute. In some implementations, a single query attribute and / or a single document attribute can itself be a combination of attributes. For example, the single document attribute can be the simultaneous occurrence of a specific first content in a specific first field of a document; and a specific second content in a specific second field of the document.
[0102] In block 458, the system generates a past interaction measure for the tuple based on a set of past interactions with documents that have one or more document features of the tuple, in response to queries with one or more query features of the tuple. In some implementations, the system can generate the past interaction measure based on a transformation of a query-document model into a "document features" and "query features" space, jointly modeled by the document feature model and the query feature model, as described here.
[0103] In block 460, the system stores the past interaction measure in association with the tuple. For example, the system can store the past interaction measure as a value for an edge connecting a query feature node, representing one or more query features of the tuple, and a document feature node, representing one or more document features of the tuple. In some implementations, the past interaction measure can be stored in a query feature-document feature model, as described herein.
[0104] In block 462, the system determines whether any remaining tuples need to be processed. If so, the system returns to block 456 to select another tuple query attribute or document attribute and performs another iteration of blocks 458, 460, and 462. The system can perform a large number of iterations of blocks 456, 458, 460, and 462 to generate a large number of past interaction measures for a large number of tuples. Such iterations can be performed sequentially and / or in parallel.
[0105] If the system determines in block 460 that there are no remaining tuples to process, the process ends in block 464. The past interaction measures generated based on procedure 400 can be used, for example, in procedure 500 described below and / or by the document measure system 130 as described herein.
[0106] Fig. Figure 5 is a flowchart illustrating an exemplary Procedure 500 for using query-dependent and / or query-independent measures of documents in response to a query to classify the documents and to provide appropriate search results based on the classification. For simplicity, the operations of the flowchart are described with reference to a system that performs the operations. This system may include various components from different computer systems, such as one or more components of the Document Measurement System 130. Although the operations of Procedure 500 are shown in a particular order, this is not intended to be restrictive. One or more operations may be rearranged, omitted, or added.
[0107] The system receives a request in block 552.
[0108] In block 554, the system identifies response documents in response to the request. These response documents may optionally contain or be restricted to access-restricted documents.
[0109] In block 556, the system identifies one or more query characteristics for the query from block 552 and identifies one or more document characteristics for a document from the response documents of block 554.
[0110] In block 558, the system generates a query-dependent measure for the document based on past interaction measures between the query characteristics and the document characteristics.
[0111] In block 560, the system generates a query-independent measure for the document based on past interaction measures in response to queries that do not have any of the query characteristics of the query received in block 552. In some implementations, the system may only execute one of blocks 558 and 560.
[0112] In block 562, the system determines whether there are any remaining documents to process. If so, the system proceeds to block 564 and identifies document characteristics for one of the remaining documents. The system then performs another iteration of blocks 558 and 560 using these document characteristics. The system can perform multiple iterations of blocks 564, 558, and 560, each time for a different response document. The system can process all response documents or a subset of them (e.g., only the top X documents based on scores generated from other ranking signals). Multiple iterations can be performed sequentially and / or in parallel.
[0113] If the system determines in block 562 that there are no remaining access-restricted documents to process, the system moves to block 566.
[0114] In Block 566, the system uses the query-dependent measures generated in multiple iterations of Block 558 and / or the query-independent measures generated in multiple iterations of Block 560 to rank the response documents identified in Block 554. For example, the system may adjust a base score for each of the response documents (e.g., a base score based on one or more other rank signals) in light of the query-dependent measure and / or the query-independent measure to produce a modified score.
[0115] In block 568, the system delivers search results for one or more of the response documents based on the classification in block 566. Delivering search results based on the classification in block 566 may include delivering search results with presentation characteristics based on the classification, such as a presentation order.
[0116] Fig. Figure 6 presents an example of the client device 106 and a display screen 140 of the computing device 106. The display screen 140 contains system interface elements 181, 182, and 183, which the user can interact with to cause the client device 106 to perform one or more actions. The display screen 140 also contains a search interface element 184, in which the user has entered a query, for example, "book order number," using a virtual keyboard or user interface input provided via a microphone. The search results 185A, 185B, and 185C are provided as a response to the query.
[0117] In Fig. 6. Search result 185A is first presented based on techniques described herein relating to the generation and use of a query-dependent measure for the document corresponding to search result 185A. Search result 185A corresponds to email 353C from Fig. 3. In some implementations, the query-dependent measure for search result 185A can be determined, at least partially, based on a past interaction measure that exists between a query attribute for the query "Book Order Number" and a document attribute based on content in the subject field and / or sender field of email 353C. In some implementations, the past interaction measures used to determine the query-dependent measure for search result 185A can be past interaction measures determined independently of email 353C. In some of these implementations, the past interaction measures can be based on different emails and / or other documents, such as emails 353A and 353B ( Fig. 3) be determined which are personal, access-restricted documents for other users and for the user who made the request in Fig. If you enter 6, they are not accessible.
[0118] It should be noted that in the example of Fig. 6. Search result 185A would not have been presented first without using the query-dependent measure. For example, using ranking signals that consider only keyword match could have caused search results 185A and 185C to be presented more prominently than search result 185A because these search results contain terms from the query in the subject lines of their corresponding emails. For example, search result 185B is for an email that has both "book" and "order" in its subject, and search result 185C is for another email that has "order" in its subject. In contrast, document 353C (the corresponding document for search result 185A) does not have either of the terms from the query in its subject line. Rather, it contains only one of the terms ("order") in the body of the email (see Fig. 3) Accordingly, the technical features described herein can be used to present search result 185A more prominently than would be the case without the technical features, and / or to present it initially, whereas it would not have been presented initially without the technical features. This can lead to the user being more likely to perceive and / or select the relevant search result 185A as the answer to the query, which can result in various technical advantages.For example, this can reduce various computing resources that would otherwise have been used if the search result had not been presented more prominently, such as resources used as a result of the following operations: the user navigates through multiple search results to locate search result 185A, the user submits a new search as a result of search result 185A not initially being noticed and / or presented, etc.
[0119] Although Fig. While the examples described in section 6 and others relate to email documents, many of the implementations described herein are additionally or alternatively applicable to other documents, such as, but not limited to, other documents explicitly described herein (e.g., media documents (e.g., image documents, audio documents, video documents), calendar entry documents, contact entry documents, other electronic communication (e.g., social media posts, chat messages)).
[0120] In situations where the systems described here collect or use personal information about users, users may be given the option to control whether programs or features collect user information (e.g., information about a user's social networks, social actions or activities, occupation, user preferences, or current geographic location), or to control whether and / or how content more relevant to the user is received from the content server. Furthermore, certain data may be processed in one or more ways before being stored or used, so that personally identifiable information is removed.For example, a user's identity can be handled in such a way that no personally identifiable information about the user can be determined, or a user's geographic location can be generalized when geographic location information (e.g., at the city, postal code, or state level) is obtained, so that a specific geographic location of the user cannot be determined. Thus, the user can control how information about them is collected and / or used.
[0121] Fig. Figure 7 is a block diagram of an exemplary computing device, Figure 710, which can optionally be used to perform one or more aspects of the techniques described herein. In some implementations, one or more of the components of Fig. 1 include one or more components of the exemplary computing device 710.
[0122] The computing device 710 typically contains at least one processor 714, which communicates with a number of peripheral devices via the bus subsystem 712. These peripheral devices may include a storage subsystem 724, which may contain, for example, a memory subsystem 725 and a file storage subsystem 726, user interface output devices 720, user interface input devices 722, and a network interface subsystem 716. The input and output devices enable user interaction with the computing device 710. The network interface subsystem 716 provides an interface to external networks and is coupled with corresponding interface devices in other computing devices.
[0123] The user interface input devices 722 can include a keyboard, pointing devices such as a mouse, trackball, touchpad or graphics tablet, a scanner, a touchscreen integrated into the display, audio input devices such as speech recognition systems, microphones and / or other types of input devices. In general, the use of the term "input device" is intended to encompass all possible device types and methods of inputting information into the computing device 710 or into a communications network.
[0124] The user interface output devices 720 can include a display subsystem, a printer, a fax machine, or non-visual displays such as audio output devices. The display subsystem can include a cathode ray tube (CRT), a flat-panel display device such as a liquid crystal display (LCD), a projection device, or any other mechanism for producing a visible image. The display subsystem can also provide a non-visual display, for example, via audio output devices. In general, the use of the term "output device" is intended to encompass all possible types of devices and methods for outputting information from the computing device 710 to the user or to another machine or computing device.
[0125] The 724 storage subsystem stores program and data constructs that provide the functionality of some or all of the modules described herein. For example, the 724 storage subsystem may contain the logic to execute selected aspects of the procedures of Fig. to execute steps 4 and / or 5.
[0126] These software modules are generally executed by the 724 processor alone or in combination with other processors. The 725 memory subsystem, used in the 724 file storage subsystem, can include several memories, including a 730 main random-access memory (main RAM) for storing instructions and data during program execution, and a 732 read-only memory (ROM) for storing fixed instructions. The 726 file storage subsystem can provide persistent memory for program and data files and can include a hard disk drive, a floppy disk drive along with associated removable media, a CD-ROM drive, an optical drive, or removable media cartridges.The modules that implement the functionality of certain implementations can be stored by the file storage subsystem 726 in the storage storage subsystem 724 or in other machines that can be accessed by the processor(s) 714.
[0127] The 712 bus subsystem provides a mechanism to enable the various components and subsystems of the 710 computing device to communicate with each other as intended. Although the 712 bus subsystem is schematically shown as a single bus, alternative implementations of the bus subsystem can use multiple buses.
[0128] The computing device 710 can be of various types, including a workstation, a server, a computer cluster, a blade server, a server farm, or another data processing system or computing device. Due to the constantly changing nature of computers and networks, the description of the in Fig. The computer device 710 shown in Figure 7 is intended only as a specific example for the purpose of illustrating some implementations. Many other configurations of the computer device 710 are possible, with more or fewer components than those shown. Fig. 7. [The text abruptly ends here, so the translation stops as well.]
[0129] According to implementations in this disclosure, methods and devices are provided that relate to using one or more document features of a document in response to a query, and optionally one or more features of the query, to determine a presentation characteristic for presenting a search result that corresponds to the document. In some implementations, measures associated with the one or more document features and / or the one or more query features can be used to determine the presentation characteristic. The measures can be based on past interactions of relevant users with other documents that share one or more of the document features with the given document, where several of the other documents are each distinct from the given document (and optionally distinct from each other).In some implementations, the document and / or other documents contain or are restricted to access-restricted documents. Besides maintaining security during the search process, these implementations offer significant reductions in network traffic for the network as a whole when scaled to a large number of users of a search engine.
[0130] Although several implementations have been described and illustrated herein, a variety of other means and / or structures may be used to perform the function and / or achieve the results and / or one or more of the advantages described herein, and any such variation and / or modification shall be considered to fall within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend on the specific application or applications for which the teachings are used. Persons skilled in the art will be able to recognize or determine many equivalents to the specific implementations described herein simply by using routine experimentation.It is understood, therefore, that the foregoing implementations are only examples and that implementations within the scope of the appended claims and equivalents thereto may be carried out differently than specifically described and claimed. Implementations of this disclosure are directed to each of the features, systems, articles, materials, equipment, and / or methods described herein. Furthermore, any combination of two or more such features, systems, articles, materials, equipment, and / or methods falls within the scope of this disclosure, provided that such features, systems, articles, materials, equipment, and / or methods are not incompatible with one another.
Claims
[1] A method that is implemented by one or more processors and includes: Receiving a request, wherein the request is entered by a user via a user interface input device of the user's computing device; identifying response documents in response to the request, wherein the response documents include access-restricted documents of the user, the access-restricted documents being accessible only to the user and any limited group of other users specified by the user; identifying one or more request attributes for the request; for each of several of the access-restricted documents: identifying one or more document attributes for the access-restricted document;and generating a query-dependent measure for the access-restricted document based on measures of past interactions between the query features and the document features, wherein each of the measures is based on a set of past interactions by corresponding users with other documents exhibiting one or more of the document features, when the other documents were presented in response to corresponding queries exhibiting one or more of the query features, and wherein the other documents include several inaccessible documents that the user cannot access; using the query-dependent measures for the access-restricted documents to determine a presentation order for the response documents;and providing one or more of the response documents in response to the presentation request based on the presentation order, the presentation being made via a user interface output device of the computing device. [2] Method according to claim 2, wherein the document features for the access-restricted document comprise a template that is contained in a specific field of the access-restricted document. [3] Method according to claim 1 or claim 2, wherein the other documents exclude one or more of the access-restricted documents. [4] Method according to any one of claims 1 to 3, wherein the other documents on which a given measure of the dimensions is based consist of inaccessible documents which the user cannot access. [5] A method according to any one of claims 1 to 4, further comprising: for each of the multiple access-restricted documents: generating a query-independent measure for the access-restricted document based on additional measures of additional past interactions with the document features in response to additional requests that do not include any of the request features, and furthermore using the query-independent measures for the access-restricted documents to determine the presentation order for the response documents.
Citation Information
Patent Citations
System for Customized Electronic Identification of Desirable Objects
AU2008261113A1
Search engine with user activity memory
US20030123443A1
Method and apparatus for determining peer groups based upon observed usage patterns
US20070150470A1
Collaborative searching by query induction
US6732088B1
Search techniques using association graphs
WO2007124430A2