Query intent specificity

By identifying and processing the intent specificity of search queries, using the intent specific machine learning model to generate intent specificity scores, it solves the problem that existing systems are difficult to deal with tail queries, and improves the accuracy and relevance of search results.

CN120216546APending Publication Date: 2025-06-27EBAY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411936757.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-27
Filing Date
2024-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing e-commerce search systems are difficult to accurately identify and process tail queries or rarely occurring queries, resulting in the search results not meeting the user's true intentions and it is difficult to find relevant information in the vast online content.

Method used

By identifying the intent specificity of the search query, generating a query vector, and using the intent specific machine learning model to train intent specificity scores to provide search results that are more in line with user intent.

Benefits of technology

Improves the accuracy and relevance of search results, reduces the load on storage components by user equipment, extends the service life of storage components, and improves the efficiency of retrieval and ranking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216546A_ABST
    Figure CN120216546A_ABST
Patent Text Reader

Abstract

Techniques described herein relate to systems, methods, and computer storage media for providing search query intent specificity. Embodiments may include identifying a search query executed using a search engine and generating a query vector for the search query by aggregating search results from the search query by a search result embedding (e.g., an item list vector). Furthermore, in some embodiments, a similarity (e.g., cosine similarity) between the query vector and the item list vector may be determined. In this way, the intention specificity of the search query can be determined. Further, in some embodiments, intent specificity may be used to train an intent-specific machine learning model for generating intent-specific scores for other search queries. Determinations may be made for accuracy and recall, among other things, based on intent-specific scores determined using the one or more trained intent-specific machine learning models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to systems, methods, and computer storage media for providing search query intent specificity. Background Art

[0002] E-commerce search software provides search results for users based on search queries and search conditions derived from the search queries, and provides a shopping experience that allows users to browse the search results. For example, an e-commerce platform can analyze and index product data and other e-commerce data so that users of the e-commerce platform can find suitable options associated with the search query. These types of software can provide various product recommendations. Summary of the Invention

[0003] At a high level, aspects described herein relate to systems, methods, and computer storage media for providing search query intent specificity, etc. In an embodiment, one or more queries (e.g., text search queries, audio search queries, etc.) can be identified. As an example, a search query can correspond to one or more items, and the one or more items can be included in an item corpus that contains a list of multiple items. By way of illustration, the item corpus can include images of items offered by sellers on an e-commerce platform (such as vehicle parts or vehicle accessories, clothing, toys, decorations, etc.). As another example, in addition to images, the item list can also include text descriptions of the items.

[0004] Based on the identification of one or more queries, a query vector can be generated for a first query by aggregating item list vectors of search results from the first query (e.g., by taking the average of the item list vectors). Based on the generated query vector, the similarity between the query vector and the item list vectors can be determined (e.g., cosine similarity, Euclidean distance, Manhattan distance, Minkowski distance, cosine similarity, Jaccard similarity, Hamming distance, correlation coefficient, another type of distance measurement, or one or more combinations thereof). Based on the determination of the similarity between the query vector and the item list vectors, the intent specificity of the search query can be determined (e.g., by taking the average of two or more similarities).

[0005] In some embodiments, a query vector can be generated by aggregating search result embeddings (e.g., by determining an average of each associated value of the search result embeddings). In some embodiments, the query vector (or item list vector, search result embedding, etc.) can be generated via one or more of the following or a combination of one or more thereof: a word2vec embedding model, a doc2vec embedding model, an image embedding model (e.g., a convolutional neural network), a generative adversarial network, a recurrent neural network, a transformer, an autoencoder, a contrastive language-image pre-training model, a Siamese neural network, another type of model. As another example, the similarity between the query vector and the item list vector can be determined based on the item list vector that has a specific distance from the query vector, wherein the distance between the search result embedding and the query vector can be measured via one or more of the Euclidean distance, Manhattan distance, Minkowski distance, cosine similarity, Jaccard similarity, Hamming distance, correlation coefficient, another type of distance measurement, or a combination of one or more thereof.

[0006] Additionally, in additional embodiments, intent specificity can be used to train an intent-specific machine learning model to generate intent-specific scores for other search queries. For example, an intent-specific machine learning model can be trained using an intent-specific distribution that is generated using search results from one or more queries that are associated with a higher number of user interactions compared to the number of user interactions with search results for other search queries. In some embodiments, multiple intent specificities can be used to train the intent-specific machine learning model, where the multiple intent specificities are generated using search results for multiple queries.

[0007] In some embodiments, the intent-specific machine learning model can include a bidirectional encoder representation from a transformer (BERT), a large language model, a generative pre-trained transformer, a text-to-text transfer transformer, a conditional transformer language model, a generative adversarial network, a vector quantization variational autoencoder, a contrastive language-image pre-training model, other types of machine learning models, or a combination of one or more thereof. As a non-limiting example, one or more BERTs can be applied to generate intent-specific scores for search queries of a specific query type (e.g., a tail query type). For example, the intent specificity determined for other query types (e.g., a head query type, a torso query type, etc.) can be used to train the intent-specific machine learning model to generate intent-specific scores for search queries of another query type (e.g., a tail query type).

[0008] Based on providing an intent - specific score for a search query (e.g., of a specific query type) using a trained intent - specific machine - learning model, one or more search results can be provided (e.g., via a user interface). In some embodiments, the one or more search results can be provided based on whether the intent - specific score for the search query (e.g., of the specific query type) is higher or lower than an intent - specific score threshold. As an example, based on determining that the intent - specific score for a search query of the specific query type is higher than the intent - specific score threshold, the one or more search results can be provided based on an indication of precision for the search query of the specific query type (rather than based on an indication of recall).

[0009] As another example, based on determining that the intent - specific score is higher than the intent - specific score threshold, the one or more search results can be provided based on an indication that the search query is a query - relevant factor. In yet another example, based on determining that the intent - specific score for a search query is lower than the intent - specific score threshold, the one or more search results can be provided based on an indication that the search query is a query - irrelevant factor.

[0010] In these ways, the techniques described herein can use the intent - specificity and intent - specific scores generated using a trained intent - specific machine - learning model for query suggestions associated with higher specificity (e.g., for automatic competitive queries, for suggesting related searches), generate refinement options and search results from refinement options with higher specificity, perform complex management of query - relevant and query - irrelevant factors, improve retrieval and ranking operations for null and low recall, etc.

[0011] This summary is intended to introduce some concepts in a simplified form that will be further described in the detailed description of the present disclosure. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to help determine the scope of the claimed subject matter. Additional objectives, advantages, and novel features of the technology will be partially described in the subsequent description, and will become partially apparent to those skilled in the art upon examination of the present disclosure or through practice of the technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The present technology will be described in detail below with reference to the accompanying drawings, in which:

[0013] Figure 1 includes an example operating environment suitable for providing intent - specificity and intent - specific scores according to embodiments described herein;

[0014] Figure 2 is a block diagram of search vector generation according to embodiments described herein;

[0015] Figure 3 A block diagram of providing an intent - specific score using a trained intent - specific machine learning model according to an embodiment described herein;

[0016] Figure 4 An example flow chart showing the generation of query vectors and intent - specificity that can be used to train an intent - specific machine learning model according to an embodiment described herein;

[0017] Figure 5 An example flow chart showing the use of a trained intent - specific machine learning model to generate an intent - specific score according to an embodiment described herein;

[0018] Figure 6 An example flow chart showing the use of a trained intent - specific machine learning model to provide intent - specificity and an intent - specific score according to an embodiment described herein; and

[0019] Figure 7 An example computing device suitable for implementing the techniques according to an embodiment described herein. Detailed Description

[0020] General query search refers to the process of using a search engine to find information through a network such as the Internet (e.g., by entering a text query). For these general search queries, a user can enter a series of words or phrases into the search system, and the search engine returns a list of items (including web pages, documents, images, or other types of files that are considered relevant to the query).

[0021] Some search queries may involve head queries, torso queries, or tail queries. For example, compared to the search results generated for tail queries, head queries and torso queries generally generate search results with higher user click - through rates or other types of user interactions of interest. As another example, compared to the search results generated for torso queries, head queries generally generate search results with higher user click - through rates or other types of user interactions of interest. In yet another example, compared to head queries and torso queries, tail queries generally have a higher volume of a small number of queries with longer words or multi - word phrases. In other words, head queries are typically queries that are frequently searched and receive most of the user engagement, torso queries are typically queries that have a significant amount of search volume and engagement but not as much as head queries, and tail queries are not often submitted and do not receive many clicks or engagement. Typically, tail queries may be more personalized and longer than head queries.

[0022] Compared to tail queries that have little historical data to rely on, frequent queries (which can be referred to as head queries) provide more click data, making it more difficult to analyze tail queries using ranking algorithms. One of the challenges faced by search engines is how to handle tail queries or queries that occur very rarely. For example, the tail of the query frequency distribution may include noisy and sparse logs of these tail queries. Although tail queries may be individually uncommon, they can also together constitute a large portion of the queries. As another example, this large portion of tail queries may exhibit a high degree of variability, which also makes it difficult (e.g., by using ranking algorithms) to analyze search results.

[0023] In terms of search queries, a user's search intent may be very specific, such that if a searcher is looking for a particular product, anything other than an exact match may be useless. At the other end of the spectrum of user intent, searchers with more open or less specific intents may expect to see other desired but less relevant search results when providing a search query. Without properly considering the specificity of the search intent, the management of the precision-recall trade-off in retrieval by a search engine and the management of the trade-off between query-related and query-unrelated factors in ranking may become less precise and less efficient.

[0024] For example, many traditional search engine systems consider the specificity of search intent inaccurately or inefficiently. Additionally, for instance, due to the query frequency distribution including noisy and sparse logs of these tail queries, the lack of information related to tail queries, or because tail queries exhibit a high degree of variability, these systems also consider the intent of users for tail queries inaccurately or inefficiently. In this way, the search results ultimately provided by these systems do not conform to the true intent of users providing a particular type of search query (e.g., tail queries). For example, traditional systems typically identify and generate search results based on the lack of information related to tail queries and, as a result, are unable to adapt to scenarios where specific search results are to be generated or where users have open-ended or less specific intents. In these ways, these traditional systems may thus fail to generate high-quality targeted results tailored to specific user preferences.

[0025] In addition, due to the quantity of available online content or stored content and the diversity of such content, enhanced and more targeted search techniques are indispensable to the technical functionality of an e-commerce search system. For example, the Internet hosts billions of web pages, images, videos, social media data, advertising data, and other forms of data. Without sophisticated search techniques and ranking algorithms that are fine-tuned for tail queries and other queries with less information about what the user is doing with those specific search results, it would be very difficult to find and process relevant information in this vast ocean of content. Additionally, prior systems have not been configured with comprehensive logic and infrastructure to effectively generate and provide reliable intent specificity for tail queries and other queries with less information about what the user is doing with those specific search results. Helping a user find a product through a query can be challenging and may require sophisticated algorithms for determining the intent of the user when providing a search query.

[0026] An e-commerce method and system are desired (e.g., for tail queries and other queries with less information about what the user is doing with those specific search results) to accurately identify accurate intent specificity. It is also desired to enhance computer network component communication between e-commerce system components and between the e-commerce system and the user's client device, where the user is using an e-commerce platform to provide a search query and the e-commerce system can identify the intent specificity for that search query. The techniques described herein achieve these goals and provide various improvements to the problems specific to the prior systems described above.

[0027] For example, the techniques discussed herein can generate search results based on intent - specific determinations made for a provided search query, such that the search results are more in line with a particular user's preferences or a particular user's intent. Additionally, rather than continuously receiving search queries provided by the user and filtering suggestions to arrive at a desired list of items, the techniques discussed herein, based on providing intent - specificity for a search query to generate search results that are more in line with a particular user's preferences or intent, can also reduce the amount of time each operating system or other component spends processing user requests and reduce excessive computer input / output operations (e.g., excessive physical read / write head movement on non - volatile disks). Further, since the user device does not have to touch the storage device to perform read or write operations for successive additional queries, added search query term descriptions, or application of various filters, the techniques described herein can also reduce the physical wear of the storage components. (For example, read / write heads are inherently very mechanical and are prone to information access errors due to the precise movements they must make to locate specific data. Such information access errors are more likely to occur when there is excessive computer I / O. Additionally, each repeated input (e.g., re - entering a search query, adding a search query term description, or applying various filters) also requires saving data to memory, thereby unnecessarily consuming storage space.)

[0028] In an embodiment of the present disclosure, a computer - implemented method can include: identifying a search query provided at a search engine. In some embodiments, the search query can have a first query type that is different from a second query type, and the first query type corresponds to a higher number of user interactions compared to the second query type. For example, the first query type of a plurality of historical queries can correspond to a higher number of user interactions with the search results provided for those historical queries of the first type. That is, the first query can be a search query previously provided by the user. In some embodiments, the first query type is a head query type and the second query type is a tail query type. In other embodiments, the first query type is a torso query type and the second query type is a tail query type.

[0029] The identified search queries provided at a search engine can correspond to a plurality of user interactions with search results of the search queries that are above a threshold. For example, in some embodiments, a second query type is a search query type that has a lesser amount of user interactions with search results of those queries for the second query type compared to a search query for a first query type. As another example, in some embodiments, the second query type corresponds to a plurality of user interactions with search results of those queries for the second query type that are below a threshold. In some embodiments, the second query type is a tail query type. In some embodiments, compared to a head query type or a torso query type, the second query type has a higher amount of a small number of queries with longer words or multi-word phrases. In some embodiments, the first query type can be a head query type or a torso query type.

[0030] The term "user" can correspond to a human, a particular entity, a robot, another particular machine, etc.

[0031] The term "user interaction" can include, for example, an item or service associated with a list of items and corresponding to a previous purchase, an item or service corresponding to a previous click (e.g., a selection of a list of items, or a selection of an image within a list of items), a ranking provided for a particular item or service, an item or service indicated as "liked", an item or service indicated as "favorited", other indications or annotations provided by a user for an item or service, scrolling within a list of items for a particular period of time, hovering over an image of an item within a list of items for a particular period of time, pausing for a particular period of time between viewing lists of items, previous search query modifications and applied filters, other types of user interactions, or one or more combinations thereof.

[0032] The term "item" as referred to herein can represent something that can be identified in response to a search query (e.g., a search query within an e-commerce platform). For example, an item can be a commodity, a software product, a tangible item, an intangible item (e.g., computer software, electronic document, movie video, song audio, electronic photo, etc.), another type of item, or one or more combinations thereof.

[0033] The term "list of items" as referred to herein can include a title, one or more images, one or more videos, metadata, item descriptions, other list of items data, or one or more combinations thereof.

[0034] The computer-implemented method may further include: generating a query vector for a first query (e.g., based on historical user interactions with search results from the first query and based on various determinations associated with search results specific to the first query). For example, intent specificity may be determined for a query set including the first query. The intent specificity for the first query may be determined based on: aggregating item list vectors of search results for the first query to generate a query vector for the first query, determining the similarity between the item list vector and the query vector (e.g., cosine similarity, a specific distance between the query vector and the item list vector, etc.), and determining the average value of the similarity.

[0035] In addition, an intent-specific machine learning model may be trained based on the intent specificity (e.g., based on query vectors generated for a head query, a torso query, or one or more combinations thereof) to generate an intent-specific score for a search query (e.g., for a query with a lower interaction history with search results of these queries). In some embodiments, the intent-specific machine learning model may include Bidirectional Encoder Representations from Transformers (BERT), large language models, Generative Pretrained Transformers, Text-to-Text Transfer Transformers, Conditional Transformer Language Models, Generative Adversarial Networks, Vector Quantized Variational Autoencoders, Contrastive Language-Image Pretraining models, other types of machine learning models, or one or more combinations thereof.

[0036] The computer-implemented method may further include: using the trained intent-specific machine learning model to provide an intent-specific score (e.g., for a second query of a second query type). In some embodiments, based on determining that the intent-specific score for the second query is higher than an intent-specific score threshold, the second query may be used as a query-related factor for generating search results for the second query. In some embodiments, based on determining that the intent-specific score for the second query is lower than the intent-specific score threshold, the second query may be used as a query-unrelated factor (e.g., for generating search results for the second query, for suggesting auto-completion for the second query, for auto-completing the second query, for suggesting related searches, etc.). In some embodiments, when managing the trade-off between precision and recall for the second query, the threshold may vary according to one or more of the intent specificity associated with the first query type or the intent-specific score associated with the second query type.

[0037] While online marketplaces that utilize the disclosed techniques to identify and retrieve items and lists of items for generating and regenerating three-dimensional (or other dimensional) environments may be referenced, it will be understood that the techniques discussed herein can be used in the context of more general online search engines (e.g., identifying and retrieving web results, providing answers in response to search queries, providing relevant searches, providing or identifying advertisements, providing other types of results, or one or more combinations thereof).

[0038] After providing some example scenarios, the techniques suitable for performing these examples will be described in more detail with reference to the accompanying drawings. It will be understood that additional systems and methods for providing improved search results and navigation can be derived from the description of the following techniques.

[0039] Now turning to Figure 1 , Figure 1 an example operating environment 100 in which embodiments of the present disclosure can be employed is shown. Specifically, Figure 1 a high-level architecture of the operating environment 100 with components according to embodiments of the present disclosure is shown. Figure 1 The components and architecture in

[0040] are intended as examples, as described at the end of the detailed description.

[0041] The intent-specific application client 102 can be a device having the ability to use a wireless communication network and can also be referred to as a "computing device", "mobile device", "client device", "user device", "wireless communication device", or "UE". In some embodiments, the user device can take various forms, such as a PC, laptop, tablet, mobile phone, PDA, server, or any other device capable of communicating with other devices using wireless communication (e.g., by sending or receiving signals). Broadly, the intent-specific application client 102 can include a computer-readable medium storing computer-executable instructions executed by at least one computer processor. An example of the intent-specific application client 102 includes the computing device 700 described herein with reference to Figure 7 .

[0042] As shown in example environment 100, the intent-specific application client 102 may be able to communicate with a search engine 110, a server 106, or a database 120 via a network 108. Other embodiments of the example environment 100 may include additional intent-specific application clients. The intent-specific application client 102 may be operated by a user (such as a person, a machine, a robot, one or more of other user device operators, or a combination of one or more thereof).

[0043] The intent-specific application client 102 may be associated with a seller interface and a buyer interface. The intent-specific application client 102 may also display image data, text data, extended reality data, other types of data, or a combination of one or more thereof based on server 106 operations and search engine 110 operations (e.g., operations associated with a query vector generator 112, an intent-specific score generator 114, an intent-specific threshold generator 116, a search result generator 118, etc., or a combination of one or more thereof).

[0044] In an embodiment, the network 108 may be a local area network (LAN), a wide area network (WAN), a mesh network, a hybrid network, or some other wired or wireless network. The network 108 may be the Internet or some other public or private network. The intent-specific application client 102 may be connected to the network 108 via a network interface (such as via wired or wireless communication). Other embodiments of the example environment 100 may include additional servers connected to the network 108.

[0045] Generally, the server 106 is a computing device that implements the functional aspects of the operating environment 100. In an embodiment, the server 106 represents a backend or server-side device. In some embodiments, the server 106 may be an edge server that receives user device requests (e.g., search queries, etc.) and coordinates the implementation of these requests (e.g., sometimes through other servers).

[0046] Additionally, the server 106 may include a computing device (e.g., Figure 7 computing device 700). Although the server 106 is logically shown as a single server, in other embodiments, the server 106 may be a distributed computing environment that encompasses multiple computing devices located at the same physical location or different physical locations. For example, in some embodiments, the server 106 may be connected to the database 120, or in other embodiments, the server 106 may communicate with multiple servers, where each server shares the database 120 or each server has its own database.

[0047] In an embodiment, the intent-specific application client 102 is a client-side or front-end device, and the server 106 is a back-end or server-side computing device. In some embodiments, the server 106 and the intent-specific application client 102 may implement functional aspects of the exemplary operating environment 100 (such as one or more functions of the search engine 110). It will be understood that some implementations of the technology will include client-side or front-end computing devices, back-end or server-side computing devices, or both, that perform any combination of functions from the search engine 110 as well as other functions or combinations of functions.

[0048] The search engine 110 may access the database 120 to perform tasks associated with the intent-specific machine learning model 122 (e.g., operations for using the intent-specific machine learning model 122 via the intent-specific score generator 114). For example, via the intent-specific application client 102 (e.g., a prompt interface, a search bar, etc.), a user may convey a request (e.g., a tail search query, or another type of search query with limited associated historical user interactions) to the search engine 110 to process the request. Based on conveying the request, the search engine 110 may perform operations (e.g., training, generating, deploying and integrating, mapping, rotating, and control operations) with components of the database 120 (such as the intent-specific machine learning model 122) to ensure processing of the request. The search engine 110 may utilize the intent-specific machine learning model 122 to create, generate, or produce content, data, or output. The search engine 110 may use the intent-specific machine learning model 122 to perform various tasks across different domains to provide improvements in automation, efficiency, and search result generation.

[0049] Generally, the search engine 110 and the intent-specific application client 102 may operate in a server-client relationship to provide search results and various suggestions to the intent-specific application client 102 based on the operations of the query vector generator 112, the intent-specific score generator 114, the intent-specific threshold generator 116, the search result generator 118, other operations of the search engine 110, or one or more combinations thereof. For example, a user may provide a search query (e.g., a tail query) from the intent-specific application client 102 to perform search engine 110 operations. Based on the search query, the search engine 110 may use the intent-specific machine learning model 122 to determine an intent-specific score for the search query. For example, the search query may be processed such that interface data (e.g., image data, text data, etc.) is generated and transmitted via the intent-specific application interface 104.

[0050] In an embodiment, based on one or more of the utilization of the item list embedding 124, the user interaction history 128, other data (e.g., other data stored in the database 120), or one or more combinations thereof, the query vector generator 112 can generate one or more queries for a specific type (e.g., for a head query, a torso query) or one or more query vectors for one or more parts of a query of a specific type. In other words, the query vector generator 112 can generate one or more query vectors for a specific type of historical search query such that the historical search query has a threshold number of prior user interaction histories for search results for each historical search query.

[0051] In some embodiments, the query vector generator 112 can generate a query vector based on the user interaction history 128 with respect to the search results for the search query. For example, search result embeddings (e.g., item list search result embedding 124, web search result embedding, video search result embedding, news article search result embedding, social media search result embedding, etc.) can be used to generate the query vector. In some embodiments, the search engine 110 can use a search result embedding having a specific distance measurement from the query vector to generate one or more intent specificities (e.g., for a head or torso query). In an embodiment, a query vector can be generated by aggregating search result item list vectors for the search results of the query, and the search result item list vectors are generated based on data within the item list (e.g., images of items in the item list, text descriptions of items in the item list, etc.). The search result item list vectors can also be generated based on the user historical interactions with each corresponding item list.

[0052] In some embodiments, the search engine 110 can determine an intent specificity to store in the intent specificity 130 for a query (e.g., a head query) based on determining the similarity between the query vector and the search results. For example, one or more average cosine similarities between the query vector for a first query and its corresponding search result item list vector can be determined. As another non-limiting example, the intent specificity (e.g., for a head query, a torso query, other types of queries having a specific number or amount of historical user interaction data for search results, or one or more combinations thereof) can be determined based on a specific distance between the query vector and its search result item list vector. For example, the specific distance can be measured via one or more of the Euclidean distance, Manhattan distance, Minkowski distance, cosine similarity, Jaccard similarity, Hamming distance, correlation coefficient, another type of distance measurement, or one or more combinations thereof.

[0053] In some embodiments, the query vector generator 112 may aggregate or concatenate the search result item list embeddings for the query from the item list embeddings 124 to generate a query vector (e.g., using one or more of a portion of these search result item list embeddings). In an embodiment, the intent specificity for the query may be determined based on the cosine similarity between the query vector generated according to the aggregation or concatenation and at least a portion of the search result vectors. For example, the intent specificity may be determined using each search result vector for the query. As another example, the intent specificity may be determined using a first portion of the search result vectors (e.g., within a first cosine similarity range) and a second portion of the search result vectors (e.g., within a second cosine similarity range). In some embodiments, based on historical user interactions associated with a particular item list and tags within that item list associated with a search query (e.g., a head query, a torso query, another type of query having a specific quantity or amount of historical user interaction data for search results), search result item list vectors may be generated or used to generate search query vectors.

[0054] In some embodiments, the query vector generator 112 may generate a query vector (e.g., an aggregation of specific item list vectors) and search result vectors (e.g., item list vectors) via one or more of the following or a combination of one or more of them: a word2vec embedding model, a doc2vec embedding model, an image embedding model (e.g., a convolutional neural network), a generative adversarial network, a recurrent neural network, a transformer, an autoencoder, a contrastive language-image pre-training model, a Siamese neural network, another type of model. For example, an image embedding model (e.g., a convolutional neural network) may be used to generate item list vectors based on processing one or more images of an item list for a particular item. As another example, item list vectors may be generated by applying both an image embedding model and a language model (e.g., a large language model) to an item list (e.g., including both images and text descriptions). In yet another example, item list vectors may be generated by applying a multimodal model (e.g., a contrastive language-image pre-training model) capable of processing both images and text.

[0055] In some embodiments, the search engine 110 may generate an intent - specific distribution of search results for a first query from a first query type (e.g., a head query) based on user interactions with the search results. The search engine 110 may also generate an intent - specific distribution of search results for multiple queries from the first query type based on user interactions with the search results, generate an intent - specific distribution of search results for a second query from a second query type (e.g., a torso query) based on user interactions with the search results, generate an intent - specific distribution of search results for multiple queries from the second query type based on user interactions with the search results, or one or more combinations of the above operations, where the first query type is different from the second query type, and where each of the two query types has a specific quantity or amount of historical user interaction data for the search results (e.g., where each of the two query types has a greater specific quantity or amount of historical user interaction data for the search results compared to the quantity or amount of historical user interaction data for the search results of a tail query type).

[0056] In some embodiments, the intent - specific distribution may be generated based on the item - list embedding 124 and the user - interaction history 128. For example, the intent - specific distribution of search results for a particular query may refer to an arrangement or layout of scores that are assigned to documents or item lists returned by a search engine (e.g., search engine 110) in response to a received search query. For example, the intent - specific distribution may be generated by using the ranking of search results that were previously used when presenting the search results via a display. In some embodiments, the intent - specific distribution may be generated based on one or more of previous clicks, cost - per - click data, click - through rate, impression share (e.g., based on potential impressions), return on ad spend, provided ratings, number of views over a particular time period, one or more other types of historical user interactions, or one or more combinations thereof. In some embodiments, the intent - specific distribution may be generated by the query vector generator 112 extracting specific text data, image data, metadata, item - list embeddings, etc. from the item list of search results for a particular historical search query.

[0057] In an embodiment, the user interaction history 128 can include user interaction data of a particular user associated with historical search queries (e.g., queries for the head or torso), user interaction data of other particular users associated with the historical search queries, other types of historical user interactions, or one or more combinations thereof. In some embodiments, one or more models (e.g., generative adversarial networks, convolutional neural networks, recurrent neural networks, transformers, autoencoders, contrastive language-image pre-training models, siamese neural networks, another type of machine learning model, or one or more combinations thereof) can be used to generate the item list embedding 124 based on the user interaction history 128 and based on the historical search queries (e.g., head queries and torso queries).

[0058] The database 120 (storing the intent-specific machine learning model 122, the item list embedding 124, the search query vector 126, the user interaction history 128, the intent specificity 130, and the intent specificity score 132) can store computer instructions (e.g., software program instructions, routines, or services) or the models (e.g., the intent-specific machine learning model 122) used in the embodiments described herein. For example, the database 120 can store computer instructions for implementing the functional aspects of the search engine 110. Although depicted as a single database component, the database 120 can be embodied as one or more databases (e.g., a distributed computing environment spanning multiple computing devices), or can be located in the cloud. In other embodiments, the intent-specific machine learning model 122 can be stored in a separate database. In an embodiment, the search engine 110 can be configured to run any number of queries on the database 120.

[0059] In an embodiment, based on a user interaction history 128 having specific user interaction history stored as a data structure (e.g., log records or statistical data in one or more tables such as a relational database table), an item list embedding 124, a search query vector 126, or an intent specificity 130 can be generated and stored. In some embodiments, one or more web caches (cookies) at a search engine 110 can be used to generate a data structure for the user interaction history 128 to organize the user interaction history by user, by search query type (e.g., head query), organize the search query type based on keywords or item list vectors, organize the search query type based on the amount of user interaction, etc., or one or more combinations thereof. In some embodiments, using specific user interactions with search results from historical queries, a data structure for the user interaction history 128 can be generated, where the historical queries are from a user device fingerprint or IP address of a user device that provided the historical queries. For example, a data structure for the user interaction history 128 can be generated using specific user interaction history from user devices within a specific geographic region, user interaction history from user devices on a specific network, user interactions during a specific time period, etc., or one or more combinations thereof.

[0060] In some embodiments, a search query vector 126 can be generated based on the item list embedding 124 or the user interaction history 128 being located within a specific location (e.g., a row or column of a table) of a data structure and based on an increment (e.g., the amount of clicks, purchases, "likes", shares via a specific digital platform, view count of an item list). In some embodiments, a search query vector 126 can be generated based on generating an item list embedding 124 according to a specific user rating (e.g., higher than 4.6 stars and more than 30 ratings). In some embodiments, one or more machine learning models (e.g., convolutional neural network, recurrent neural network, deep learning model, gradient boosting decision tree, random forest decision tree, another type of decision tree, generative adversarial neural network, regression model, another type of machine learning model, or one or more combinations thereof) can be used to generate the item list embedding 124. In some embodiments, the item list embedding 124 can be generated by cascading text embeddings and image embeddings for an image and text description of an item list.

[0061] In some embodiments, the search engine 110 can perform tasks concurrently. For example, the query vector generator 112 can concurrently generate query vectors for multiple head queries and query vectors for multiple torso queries for storage at the search query vector 126. As another example, the search engine 110 can concurrently generate intent specificities 130 for head queries and for torso queries (e.g., intent specificities indicating the relevance of an item or list of items in the context of the current query without regard to the user's desired intent), or can concurrently generate intent specificities 130 for multiple head queries. In some embodiments, for head queries and torso queries, the search engine 110 can generate an intent specificity distribution based on the item list embedding 124 and the user interaction history 128. In some embodiments, the intent specificity distribution can be generated based on the item list embedding 124 or the user interaction history 128 being located within a particular location (e.g., a row or column of a table) of a data structure and based on increments (e.g., the amount of clicks, purchases, "likes", shares via a particular digital platform, view counts of the item list).

[0062] In some embodiments, a specific weight can be applied to a specific search result (e.g., for a head query or a torso query) to generate the search query vector 126 via the query vector generator 112. For example, a specific item list from a search result can be given a higher weight based on a specific user interaction history 128 (e.g., a specific number of clicks, a specific number of purchases, a specific number of views, etc.). By way of illustration, the item list for a first search result can be given a higher weight than a second search result, where that item list has a combination of user interaction history (e.g., both purchases and views) involving more user interactions compared to another item list for the second search result. In some embodiments, when training the intent specific machine learning model 122 based on the search query vectors 126 generated for head queries and torso queries, in some embodiments, the search query vectors 126 associated with a greater number of user interaction histories or the search query vectors 126 associated with a higher intent specificity can be given a higher weight.

[0063] In an example embodiment, the database 120 or its components can include a search index (e.g., an inverted index) that can be used by the search engine 110, but other index forms are possible. For example, the database 120 or its components can include one or more tables within the search index for the search query vector 126. In another example, the particular location of the item list embedding 124 can be based on a particular color of an item, the size or dimensions of an item, the intent specificity 130, other identifiers associated with the item, or one or more combinations thereof. In other embodiments, the components of the database 120 can be discrete components separate from the search engine 110, or can be incorporated or integrated into the search engine 110.

[0064] In an embodiment, the intent-specific machine learning model 122 can include a natural language understanding model, a large language model, a text-to-speech engine, an automatic speech recognition engine (e.g., a recurrent neural network, another transformer model, or another type of machine learning technique that can perform automatic speech recognition), Bidirectional Encoder Representations from Transformers (BERT), Embeddings from Language Models (ELMo), Bidirectional Long Short-Term Memory Networks (BiLSTM), etc., or one or more combinations thereof. In some embodiments, the intent-specific machine learning model 122 can include a first larger model and a second smaller model, where the larger model has a greater number of transformer layers than the smaller model. In some embodiments, the intent-specific machine learning model 122 can include a first model whose hidden size of a single transformer layer is larger than the hidden size of a single transformer layer of a second model.

[0065] In some embodiments, the intent-specific machine learning model 122 can be trained using offline head query vector aggregation and its associated search results, offline torso query vector aggregation and its associated search results, offline browse node aggregation and its associated search results, or a combination of one or more thereof. Each vector aggregation can be generated (e.g., by the query vector generator 112) based on the user interaction history 128 associated with the associated search results. In some embodiments, the intent-specific machine learning model 122 can be trained based on query vectors generated using one or more bag-of-words models. For example, the average of clicks or purchased item lists from search results for a head query, along with BERT or another transformer, can be used to generate a head query type vector. As another example, the average of clicks or purchased item lists from search results for a torso query, along with BERT or another transformer, can be used to generate a torso query type vector.

[0066] In some embodiments, the intent-specific machine learning model 122 can be trained based on determining an average of cosines between a head query and item list embeddings of search results for that head query. Additionally or alternatively, in some embodiments, the intent-specific machine learning model 122 can be trained based on determining an average of cosines between a torso query and item list embeddings of search results for that torso query. In some embodiments, a head query is a more specific query than a torso query, such that the average of cosines between the head query and those associated item list embeddings has more closely clustered item list embeddings closer to that average compared to the cosine average associated item list embeddings for the torso query. In some embodiments, the item list embeddings generated for head and torso queries are generated without considering the category classification of the corresponding item lists.

[0067] An intent - specific machine - learning model 122 can be trained via an intent - specific score generator 114 to generate an intent - specific score 132 for search queries of different query types. For example, the intent - specific machine - learning model 122 can be trained to generate an intent - specific score 132 for a search query type that has a lower amount of associated historical user interaction with search results compared to the query type used to train the intent - specific machine - learning model 122. For example, search query vectors 126 generated for head queries and torso queries can be used to train the intent - specific machine - learning model 122 to generate an intent - specific score 132 for tail search queries. In some embodiments, an offline browsing node can aggregate its associated search results for head browsing nodes and torso browsing nodes to train the intent - specific machine - learning model 122 to generate an intent - specific score 132 for a tail browsing node (e.g., a browsing node that has a lower amount of associated historical user interaction with search results compared to other browsing nodes from which the intent - specific machine - learning model 122 is trained).

[0068] In some embodiments, the intent - specific machine - learning model 122 can include BERT, and a single output layer and the mean squared error (MSE) loss function of BERT can be used to generate an intent - specific score 132 for tail search queries. In this way, the intent - specific score generator 114 can identify patterns mined from head and torso queries to generate an intent - specific score 132 for tail search queries. In some embodiments, the intent - specific machine - learning model 122 can include BERT, and a single output layer and the mean squared error (MSE) loss function of BERT can be used to generate an intent - specific score 132 for browsing nodes with a lower amount of associated historical user interaction. In this way, the intent - specific score generator 114 can identify patterns mined from browsing nodes with a higher amount of associated historical user interaction to generate an intent - specific score 132 for browsing nodes with a lower amount of associated historical user interaction.

[0069] In some embodiments, the intent-specific machine learning model 122 can be retrained based on feedback provided by the user (e.g., modifying the query vector based on usage feedback, based on additional user interactions stored at the user interaction history 128, updating the intent-specific score threshold via the intent-specific threshold generator 116 based on usage feedback). For example, in some embodiments, in response to receiving a tail query, the query vector can be narrowed or expanded based on user feedback when providing search results via the search result generator 118. As another example, the updated user interaction history 128 can be used to modify the query vector (e.g., when the user makes a purchase after applying a specific filter when receiving search results from the search result generator 118 in response to providing a tail query).

[0070] In yet another example, based on the provided feedback (e.g., the user enters another search query without selecting any of the provided search results), the intent-specific score threshold can be updated via the intent-specific threshold generator 116 for the modified query vector. For example, if the updated specificity score is high, the intent-specific score threshold can be adjusted (e.g., adjusted based on the average of the cosines between the search results of the head query) by the intent-specific threshold generator 116. In an embodiment, instead of using a single cosine threshold for each tail query provided to the intent-specific machine learning model 122, the threshold can be changed according to the percentile of the intent specificity used when training the intent-specific machine learning model 122 to determine the intent-specific score. In this way, using the intent-specific threshold generator 116 to generate intent-specific score thresholds using different percentile values can further allow the search result generator 118 to determine a specific trade-off metric at runtime for tail search queries.

[0071] The search result generator 118 can provide search results (e.g., for a tail query) based on leveraging the intent-specific score generator 114 and the intent-specific machine learning model 122. In some embodiments, the search result generator 118 can provide an intent-specific score 132 for a received tail query or another query type based on using the trained intent-specific machine learning model 122, where the other query type is associated with a specific amount of user interaction with previous search results for that other query type. In an embodiment, the intent-specific score can be used to identify other queries similar to the provided search query (e.g., other types of queries or other query terms), manage search queries with similar mean vectors but different variances, determine a range of relevance, incorporate into a relevance model to determine the relevance and desirability trade-off for generating search results, apply as a filter in other analyses, and other search engine functions.

[0072] In some embodiments, based on an intent specificity score (e.g., based on the intent specificity score being higher than an intent specificity score threshold determined by the intent specificity threshold generator 116), specific search results for a tail query (for retrieving search results using the tail query) can be indicated as exact rather than for recall, and the search terms of the tail query can be indicated as exact rather than for recall. In some embodiments, based on an intent specificity score (e.g., based on the intent specificity score being lower than an intent specificity score threshold determined by the intent specificity threshold generator 116), specific search results for a tail query (for retrieving search results using the tail query) can be indicated for recall rather than as being exact, and the search terms of the tail query can be indicated for recall rather than as being exact. In this way, the intent specificity score can be used to determine whether to truncate retrieval based on the number of search results, determine whether to truncate retrieval based on an absolute cosine value, or one or more combinations thereof. In this way, tail queries with high specificity can favor precision, while tail queries with lower specificity can favor recall.

[0073] In some embodiments, based on an intent specificity score (e.g., based on the intent specificity score being higher than or lower than an intent specificity score threshold determined by the intent specificity threshold generator 116), specific search results for a tail query can be ranked. For example, based on the intent specificity score being higher than the intent specificity score threshold, the search results can be ranked based on indicating the tail query as a query-related factor. In this way, based on the ranking and based on indicating the tail query as a query-related factor, the search results can be provided to the intent specificity application client 102 in a specific order. In another example, based on the intent specificity score being lower than the intent specificity score threshold, the search results can be ranked based on indicating the tail query as a query-unrelated factor. In this way, the ranking system of the search result generator 118 can learn the query-specific trade-off between query-related factors and query-unrelated factors for a tail query.

[0074] Figure 2 An example block diagram 200 of search vector generation for training an intent specificity machine learning model is depicted. The example block diagram 200 includes example operations for a query vector generator 202 (e.g., operations 202A, 202B, 202C, and 202D), example operations associated with a user interaction history 204 (e.g., operations 204A, 204B, and 204C), and example operations associated with an item list embedding 206 (e.g., operations 206A, 206B, and 206C).

[0075] The query vector generator 202 can analyze a search query 202A (e.g., a text search query, an audio search query, an image search query, etc., or one or more combinations thereof). For example, the query vector generator 202 can use keyword matching (e.g., matching keywords with documents, web pages, user interaction history databases, item list indexes, user profile data, etc.), natural language processing (e.g., for semantic understanding of the search query, by parsing the query, identifying one or more entities associated with the provided query or the computing device providing the query, word relation extraction or entity relation extraction, etc.), natural language understanding, query expansion, relevance identification, personalization (e.g., associated with a user profile or with the location of the computing device providing the query), context analysis, or one or more of the like to analyze the search query (e.g., a head query, a torso query).

[0076] The query vector generator 202 can also analyze search results 202B for the search query. For example, the query vector generator 202 can analyze the search results by generating an intent-specific distribution for the search results from a first query (e.g., a head or torso query). The intent-specific distribution can be generated based on item list embeddings 206 of the item list of the search results for the first query. By way of example, the intent-specific distribution for these search results can refer to an arrangement or layout of scores, data extracted from the documents or links of the item list, etc. In some embodiments, the intent-specificity can indicate the relevance of the item list or search result items in the context of the first query, regardless of user expectations or regardless of the category taxonomy of the item list.

[0077] The item list embedding 206 (206A) can be generated based on user interaction history. In some embodiments, the item list embedding can be generated based on a specific amount of user interaction (e.g., purchases, click-through rates, etc.). The item list embedding (206B) can be generated based on previous search queries and query types. For example, the user interaction history for a specific head query can be used to generate the item list embedding for one or more item lists. In some embodiments, the item list embedding (206C) can be generated based on intent-specific scores generated using an intent-specific machine learning model for a tail query. For example, a specific item list can have an item list embedding generated using the intent-specific scores for a tail query and based on the historical search results for that tail query.

[0078] In some embodiments, the query vector generator 202 may analyze search results (202B) for a search query using one or more of a generative pre-trained transformer, variational autoencoder, recurrent neural network, long short-term memory network, attention-based model, sequence-to-sequence model, conditional generative model, neural machine translation model, BERT, zero-shot learning model, another model, or one or more combinations thereof. For example, a variational autoencoder may be used to improve an image of an item list or extract relevant data from an image within the item list. As another example, a recurrent neural network may be used to identify sequences associated with identified features of an image within the item list.

[0079] The query vector generator 202 may also analyze the user interaction history (202C) of one or more search results for a first query. For example, a particular search result having a threshold number of previous user interactions may be analyzed. In some embodiments, a particular search result having a particular user device fingerprint, a particular IP address of the user device, or tracked via a particular network cache may be analyzed. In some embodiments, a particular search result having a particular item list embedding associated with the user interaction may be analyzed. In some embodiments, a particular search result having a particular number of previous clicks, cost-per-click data, click-through rate, impression share (e.g., based on potential impressions), return on ad spend, provided ratings, number of views over a particular time period, other types of historical user interactions, or one or more combinations thereof may be analyzed.

[0080] In some embodiments, the query vector generator 202 may analyze the user interaction history (202C) to identify particular search results for a particular query type (e.g., head query type, torso query type, tail query type, another query type, etc.). That is, the user interaction history 204 may be analyzed to determine the search query type (204A). In some embodiments, the user interaction history 204 of search results for a first query type (e.g., head query type) (204B) may be analyzed, and the user interaction history 204 of search results for a second query type (e.g., torso query type) (204C) may be analyzed. In some embodiments, the search results for the first query and the search results for the second query type may be analyzed based on a particular category, particular item features, price range, particular seller, particular brand, etc., described in the item description of the corresponding item list, or one or more combinations thereof. For example, the item list may have a particular label indicating a particular user interaction history or amount of particular user interaction history, as well as a label for a particular price range.

[0081] The query vector generator 202 can generate a search query vector (202D) based on one or more analyses of the search query (202A), one or more analyzed search results of the search query (202B), one or more analyses of the user interaction history with the search results (202C), or one or more combinations thereof. For example, the query vector generator 202 can generate a search query vector (202D) based on one or more aggregations of the item list embeddings 206 for the search results for the query. As another example, the query vector generator 202 can generate another query vector based on the item list embedding corresponding to the other query vector or one or more of its parts. In some embodiments, the query vector generator 202 can generate a query vector by averaging the corresponding item list embedding or one or more of its parts.

[0082] Figure 3 An example block diagram 300 for generating an intent-specific score via an intent-specific score generator 302 is shown. The example block diagram 300 includes example operations for the intent-specific score generator 302 (e.g., operations 302A, 302B, and 302C), example operations associated with the intent-specific machine learning model 304 (e.g., operations 304A, 304B, and 304C), and example operations associated with the search query vector 306 (e.g., operations 306A, 306B, and 306C).

[0083] The intent-specific score generator 302 can generate an intent-specific score (302A) for a first query type (e.g., a tail query) based on the search query vectors 306 of other types of search queries that are generated (e.g., a head query and a torso query). For example, based on generating the search query vectors 306 of other query types (e.g., a head query and a torso query), the intent-specific machine learning model 304 can be trained to generate an intent-specific score for the first query type (e.g., for a tail query). In some embodiments, the intent-specific score generator 302 can generate an intent-specific score for a browsing node associated with a specific level of user interaction history based on the search query vectors associated with other types of browsing nodes used to train the intent-specific machine learning model 304 (e.g., the intent-specific score generator 302 can generate an intent-specific score for another query type (302B)). Additionally, the intent-specific score generator 302 can generate an intent-specific score (e.g., for a tail query) based on feedback (302C). For example, a user can provide feedback via a prompt regarding the search results provided for a tail query (based on the intent-specific score).

[0084] In an embodiment, the intent - specific machine - learning model 304 can be trained based on generating (e.g., for multiple head queries) a first set of search - query vectors (304A), based on generating (e.g., for multiple torso queries) a second set of search - query vectors (304B), based on feedback provided in response to the generated intent - specific scores (304C), or one or more combinations thereof. In some embodiments, search - query vectors can be generated based on aggregating search - result embeddings of first queries of a first query type and based on user interactions with those search results (306A). As another example, the intent - specific machine - learning model 304 can be trained based on aggregating search - result embeddings of second queries of a first query type and based on user interactions with those search results. In yet another example, the intent - specific machine - learning model 304 can be trained based on aggregating search - result embeddings of third queries of a second query type and based on user interactions with those search results.

[0085] In some embodiments, search - query vectors can be utilized to generate an intent - specific distribution of search results for a first query of a first query type (306B). In some embodiments, another search - query vector can be utilized to generate an intent - specific distribution of search results for a second query of a first query type, and the generation is based on user interactions with those search results. In some embodiments, another search - query vector can be utilized to generate an intent - specific distribution of search results for a third query of a second query type, and the generation is based on user interactions with those search results. In some embodiments, search - query vectors (e.g., for head queries or torso queries) can be utilized to determine the cosine angle between a search query and an item - list vector of search results for that query (306C). In some embodiments, the search - query vector can be utilized to identify a specific range of cosine angles for determining a threshold for the specificity score.

[0086] Figure 4FIG. 400 is an example flow diagram showing query vector 408 and intent specificity 412 for generating, e.g., an intent-specific machine learning model for training to determine an intent-specific score for another query type. At 402, a search bar may receive a search query (e.g., a head and torso search query provided historically), and at 404, search results are generated. An item list embedding may be extracted or generated for each search result at 404 at 406. An average of the item list embeddings may be determined to generate query vector 408 for the search query received at 402. A cosine angle 410 between query vector 408 and each item list embedding 406 is determined. Additionally, an average of cosine angles 410 is determined as intent specificity 412 for the search query received at 402 (e.g., a head query or a torso query).

[0087] In some embodiments, a lower cosine value away from the value "1" may indicate that the intent specificity is less similar to the average (e.g., indicating an item or item list associated with an item list embedding 406 having a lower correlation in the context of the query received at 402 without considering user expectations). In some embodiments, a higher cosine value closer to the value "1" may indicate that the intent specificity is similar to the average (e.g., indicating an item or item list associated with an item list embedding 406 having a higher correlation in the context of the query received at 402 without considering user expectations). In some embodiments, another statistical aggregation other than the average may be used to generate query vector 408 or intent specificity 412 (e.g., using the 90th percentile).

[0088] Figure 5 FIG. 500 is an example flow diagram showing generating an intent-specific score 514 based on utilizing an intent-specific machine learning model. For example, an intent-specific client device 502 may generate a classification token (CLS) 504A and one or more of device tokens 504B, 504C, 504D, and 504E and provide them to one or more language transformers 506. In an embodiment, tokens 504A - 504E may be generated for a tail type query. In some embodiments, the one or more language transformers 506 include bidirectional encoder representations from transformers (BERT). In an embodiment, one or more of device tokens 504B, 504C, 504D, and 504E are identifiers for push notifications or another type of user device interaction.

[0089] Based on providing classification token (CLS) 504A and one or more of device tokens 504B, 504C, 504D, and 504E to the one or more language transformers 506, the pooling layer 508 can use one or more of the tokens 504A - 504E (e.g., using the tokens that have passed through multiple layers of the transformer encoder) to generate context embeddings for the one or more tokens 504A - 504E. For example, the pooling layer 508 can perform one or more of mean pooling (averaging the output embeddings for all tokens in the sequence to provide a global context representation for further classification) or max pooling (using the maximum value for each dimension across the output embeddings of all tokens).

[0090] After the pooling layer 508, a dropout 510 regularization technique for the neural network architecture can be performed. For example, dropout 510 can include randomly setting a portion of the input units to zero during training (e.g., training based on the intent specificity for head and torso type queries) to prevent overfitting, thereby enhancing the intent - specific machine learning model. As another example, the dropout 510 technique can be used to minimize the mean squared error loss. In another example, the training can include using specific scores generated for head and torso type queries as regression targets. As an example, the portion of units discarded via the dropout 510 technique can be a tunable hyperparameter. In embodiments associated with inference or evaluation, dropout can be turned off, and the full model (without dropout) can be used for prediction.

[0091] A linear layer 512 can be applied. For example, the linear transformation applied at the linear layer 512 can be applied to the output embeddings from the BERT model. In an embodiment, the linear layer 512 can be used for a specific downstream task (such as text classification, named entity recognition, sentiment analysis, or another downstream task). As a result, an intent - specific score 514 can be determined based on leveraging the intent - specific machine learning model.

[0092] Example flowchart

[0093] Figure 6The flowchart 600 begins with step 602 of identifying a search query provided at a search engine. For example, the identified search query can be a first query of a first query type that is different from a second query type. For example, compared to the second query type, the first query type can correspond to a higher number of user interactions. In some embodiments, the first query type is a head query type, and the second query type is a tail query type. In some embodiments, a third query type is also identified, which is different from the second query type and the first query type. For example, the third query type can correspond to a higher number of user interactions compared to the second query type and a lower number of user interactions compared to the first query type. In some embodiments, the third query type is a torso query type. In some embodiments, compared to a search query of the second query type, the first query type corresponds to a search query having fewer search terms.

[0094] At step 604, a query vector for the search query is generated by aggregating item list vectors of search results from the search query. In some embodiments, the search results have a user interaction history generated by a user device that provided the search query to the search engine. In some embodiments, the query vector is based on using Figure 1 the item list embedding 124 of Figure 1 the user interaction history 128, other data (e.g., other data stored in Figure 1 the database 120 of Figure 4 one or more of or a combination of one or more thereof. In some embodiments, the query vector is generated based on averaging

[0095] the item list embedding 406 of Figure 4 At step 606, the similarity between the query vector and the item list vector can be determined. In some embodiments, the cosine angle 410 between the query vector 408 and Figure 4 each item list embedding 406 of

[0096] At step 608, an intent-specific machine learning model is trained to generate intent-specific scores for other search queries (e.g., other types of search queries). In some embodiments, the intent-specific machine learning model is trained using the intent specificity for the first search query determined from the query vector. In some embodiments, the intent-specific machine learning model is trained based on the intent specificity distribution of the search results from the first query based on user interactions with those search results. In some embodiments, the intent-specific machine learning model is trained using multiple intent specificities of a first set of queries for the first query type. In some embodiments, the intent-specific machine learning model is trained using multiple intent specificities of a first set of queries (e.g., a head query set) for the first query type and a second set of queries (e.g., a torso query set) for a third query type.

[0097] In some embodiments, the trained intent-specific machine learning model is used to provide an intent-specific score for a specific query of the second query type. For example, the intent-specific score may be provided based on receiving a specific query of the second query type at the search engine. As a result, search results may be provided and generated for the specific query of the second query type based on the intent-specific score. For example, the intent-specific score may be used for query suggestions associated with higher specificity (e.g., for auto-competing queries, for suggesting related searches), for refining option suggestions, for managing query-related and query-unrelated factors to generate search results, and for improving retrieval and ranking operations for generating search results.

[0098] Example computing device

[0099] After an overview of embodiments of the present technology has been described, an example operating environment in which embodiments of the present technology may be implemented will be described below to provide a general context for the various aspects. First, reference is made to Figure 7 , specifically, an example operating environment for implementing embodiments of the present technology is shown and is generally designated as computing device 700. However, computing device 700 is one example of a suitable computing environment and is not intended to imply any limitation as to the scope of use or functionality of the technology. Computing device 700 should also not be construed as having any dependency or requirement related to any one or combination of the components shown.

[0100] The techniques of the present disclosure may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions (e.g., program modules) executed by a computer or other machine (e.g., a personal data assistant or other handheld device). Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The techniques may be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The techniques may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network.

[0101] Referring Figure 7 , computing device 700 includes a bus 702 that directly or indirectly couples the following devices: a memory 704, one or more processors 706, one or more presentation components 708, an input / output port 710, an input / output component 712, and an illustrative power supply 714. Bus 702 represents one or more buses (e.g., an address bus, a data bus, or a combination thereof). Although the various blocks are shown as lines for clarity, in reality, the division of the various components is not so distinct, and metaphorically, the lines would more accurately be gray and fuzzy. For example, one could consider a presentation component such as a display device to be an I / O component. As another example, a processor may also have memory. This is the nature of the technique, and again, Figure 7 the figures shown are only examples of computing devices that may be used in conjunction with one or more embodiments of the present technique. There is no distinction between such categories as "workstation", "server", "laptop computer", "handheld device", etc., because all of these categories are within Figure 7 the scope of and relate to a "computing device". Figure 7

[0102] Computing device 700 typically includes a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by computing device 700 and includes volatile and nonvolatile media as well as removable and non-removable media. By way of example and not limitation, computer-readable media may include computer storage media and communication media.

[0103] Computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to: RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by a computing device 700. Computer storage media does not itself include a signal.

[0104] Communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and includes any information delivery media. The term "modulated data signal" refers to a signal that has one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared and other wireless media. Combinations of any of the above should also be included within the scope of computer readable media.

[0105] Memory 704 includes computer storage media in the form of volatile or non-volatile memory. Memory 704 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid state memory, hard disk drives, optical disk drives, etc. The computing device 700 includes one or more processors that read data from various entities such as memory 704 or I / O components 712. The presentation component 708 presents data indications to a user or other device. Examples of presentation components include display devices, speakers, printing components, vibration components, etc.

[0106] The I / O port 710 allows the computing device 700 to be logically coupled to other devices including the I / O components 712, some of which may be built-in. Illustrative components include microphones, joysticks, gamepads, satellite dishes, scanners, printers, wireless devices, etc.

[0107] The above embodiments can be combined with one or more of the specifically described alternatives. Specifically, the claimed embodiments can include references to more than one other embodiment in the alternatives. The claimed embodiments can specify additional limitations of the claimed subject matter.

[0108] The subject matter of the present technology is described specifically herein to meet statutory requirements. However, the description itself is not intended to limit the scope of the disclosure. On the contrary, the inventors have contemplated that the claimed or disclosed subject matter may also be embodied in other ways in combination with other existing or future technologies, to include different steps or combinations of steps similar to those described in this document. Additionally, although the terms "step" or "block" may be used herein to denote different elements of the methods employed, these terms should not be construed as implying any particular order among or between the steps disclosed herein, unless and except when the order of the individual steps is explicitly stated.

[0109] For the purposes of this disclosure, the words "comprising" or "having" have the same broad meaning as the word "including", and the word "access" includes "receiving", "referencing", or "retrieving". Additionally, the word "communicate" has the same broad meaning as the words "receive" or "send" as implemented by a software or hardware bus, receiver, or transmitter using a communication medium.

[0110] Furthermore, unless otherwise indicated to the contrary, words such as "a" and "an" include both the plural and the singular. Thus, for example, when there is one or more features, the constraint of "a feature" is satisfied.

[0111] Additionally, the term "or" includes conjunctive, disjunctive, and both (thus, a or b includes a or b, as well as a and b).

[0112] For the purposes discussed in detail above, embodiments of the present technology are described with reference to a distributed computing environment; however, the distributed computing environment depicted herein is merely an example. Components may be configured to perform the novel aspects of the embodiments, where the term "configured to" or "configured" may refer to "programmed to" perform a specific task or implement a specific abstract data type using code. Additionally, although embodiments of the present technology may generally be referenced to a distributed data object management system and the diagrams described, it is understood that the described technology may be extended to other implementation contexts.

[0113] In view of the foregoing, it will be seen that the technology is well suited to achieve all of the above objects and goals, including other advantages that are obvious or inherent to the structure. It should be understood that certain features and subcombinations are useful and may be used without reference to other features and subcombinations. This is contemplated by the claims and is within the scope of the claims. Since many possible embodiments of the described technology may be made without departing from the scope, it should be understood that all of the content described or shown in the figures herein should be construed as illustrative and not restrictive.

[0114] Some example aspects of the technology that can be practiced from the foregoing disclosure include the following aspects:

[0115] Aspect 1: A computer-implemented method includes: identifying a search query executed using a search engine; generating a query vector for the search query by aggregating item list vectors of search results from the search query; determining a similarity between the query vector and the item list vectors; and generating an intent specificity of the search query based on aggregating the similarity between the query vector and the item list vectors.

[0116] Aspect 2: According to Aspect 1, wherein the search results have a user interaction history generated by a user device that provided the search query to the search engine.

[0117] Aspect 3: According to Aspect 1 or 2, wherein aggregating the item list vectors to generate the query vector includes: determining an average value of the item list vectors.

[0118] Aspect 4: According to Aspect 1, 2, or 3, wherein the similarity between the query vector and the item list vectors is determined by using Euclidean distance.

[0119] Aspect 5: According to Aspect 1, 2, 3, or 4, wherein the similarity between the query vector and the item list vectors is determined by using cosine similarity between the query vector and the item list vectors.

[0120] Aspect 6: According to Aspect 1, 2, 3, 4, or 5, wherein aggregating the similarity between the query vector and the item list vectors includes: determining an average value of the cosine similarity.

[0121] Aspect 7: According to Aspect 1, 2, 3, 4, 5, or 6, further includes: identifying a second search query executed using the search engine; generating at least one query vector for the second search query by aggregating item list vectors of search results from the second search query; determining a similarity between the at least one query vector for the second search query and the item list vectors of search results from the second search query; and generating a second intent specificity of the second search query based on aggregating the similarity between the at least one query vector for the second search query and the item list vectors of search results from the second search query.

[0122] Aspect 8: According to aspect 1, 2, 3, 4, 5, 6 or 7, further comprising: using the intent specificity and the second intent specificity to train an intent-specific machine learning model to generate an intent-specific score for an additional search query; providing a third query to the trained intent-specific machine learning model; and using the trained intent-specific machine learning model to provide an intent-specific score for the third query.

[0123] Aspect 9: A computer system includes: a processor; and a computer storage medium storing computer-usable instructions that, when used by the processor, cause the computer system to perform operations including: identifying a first set of queries having a determined intent specificity; using the intent specificity for each query in the first set of queries to train an intent-specific machine learning model to generate an intent-specific score for an additional search query; providing a second query to the trained intent-specific machine learning model; and using the trained intent-specific machine learning model to provide an intent-specific score for the second query.

[0124] Aspect 10: According to aspect 9, wherein the intent specificity of the first set of queries is determined by: generating an item list vector for each set of search results corresponding to each query in the first set of queries; using the item list vector of the search results for at least one query in the first set of queries to generate a query vector for the at least one query in the first set of queries; and determining an average of the cosine similarities between the query vector and the item list vector of the search results for the at least one query in the first set of queries to determine the intent specificity for the at least one query in the first set of queries.

[0125] Aspect 11: According to aspect 9 or 10, wherein each of the first set of queries has a prior user interaction history with the search results for each query in the first set of queries that is above a threshold.

[0126] Aspect 12: According to aspect 9, 10 or 11, wherein the second query provided to the trained intent-specific machine learning model has a prior user interaction history with the search results for the second query that is below the threshold.

[0127] Aspect 13: According to aspect 9, 10, 11, or 12, further comprising: identifying a second set of queries having a determined intent specificity, the first set of queries having a higher number of user interactions with search results compared to the second set of queries, the second set of queries having a previous user interaction history above the threshold; training the intent-specific machine learning model using the intent specificity for each query in the second set of queries; and providing an intent specificity score for the second query based on training the intent-specific machine learning model using the intent specificity for each query in the second set of queries.

[0128] Aspect 14: According to aspect 9, 10, 11, 12, or 13, further comprising: determining that the intent specificity score for the second query is below an intent specificity score threshold; determining that the second query is a query-unrelated factor based on determining that the intent specificity score for the second query is below the intent specificity score threshold; ranking the search results for the second query based on determining that the second query is a query-unrelated factor; and providing the search results for the second query based on the ranking.

[0129] Aspect 15: One or more non-transitory computer storage media storing computer-usable instructions that, when used by a user device, cause the user device to perform operations, the operations including: identifying a search query executed using a search engine; generating a query vector for the search query by aggregating item list vectors of search results from the search query; determining a similarity between the query vector and the item list vectors; and generating an intent specificity of the search query based on aggregating the similarity between the query vector and the item list vectors.

[0130] Aspect 16: According to aspect 15, further comprising: identifying a second query executed using the search engine; aggregating search result embeddings of search results for the second query to generate a second query vector; determining a cosine angle between the second query vector and each search result embedding of the search results for the second query to generate an intent specificity for the second query.

[0131] Aspect 17: According to aspect 15 or 16, further comprising: training an intent-specific machine learning model using the intent specificity for the search query and the intent specificity for the second query; providing a third query to the trained intent-specific machine learning model, the third query having a lower user interaction history for the search results of the third query compared to the first query and the second query, the intent-specific machine learning model being trained to: generate an intent-specific score for a query having a lower user interaction history for search results than the user interaction history for the first query and the second query; and providing an intent-specific score for the third query using the trained intent-specific machine learning model.

[0132] Aspect 18: According to aspect 15, 16 or 17, wherein the search result embedding includes an item list embedding.

[0133] Aspect 19: According to aspect 15, 16, 17 or 18, further comprising: determining that the intent-specific score for the third query is higher than an intent-specific score threshold; and based on determining that the intent-specific score for the third query is higher than the intent-specific score threshold, providing an indication that the retrieval of the search results using the third query is precise, rather than providing an indication for recall.

[0134] Aspect 20: According to aspect 15, 16, 17, 18 or 19, further comprising: determining that the intent-specific score for the third query is higher than an intent-specific score threshold; based on determining that the intent-specific score for the third query is higher than the intent-specific score threshold, determining that the third query is a query-related factor; based on determining that the third query is a query-related factor, ranking the search results for the third query; and providing the search results for the third query based on the ranking.

[0135] Aspect 21: A computer-implemented method includes: identifying a query of a first query type that is different from a second query type, the first query type corresponding to a higher number of user interactions compared to the second query type; generating an intent-specific distribution for search results corresponding to queries of the first query type based on user interactions with the search results; using the intent-specific distribution to train an intent-specific machine learning model to generate an intent-specific score for search queries of the second query type; providing a search query of the second query type to the trained intent-specific machine learning model; and using the trained intent-specific machine learning model to provide an intent-specific score for the search query.

[0136] Aspect 22: According to aspect 21, wherein the first query type is a head query type and the second query type is a tail query type.

[0137] Aspect 23: According to aspect 21 or 22, further comprising: determining that an intent specificity score for the search query is higher than an intent specificity score threshold; and based on determining that the intent specificity score for the search query is higher than the intent specificity score threshold, providing a set of search results for the search query based on indicating that the retrieval of the set of search results using the search query is precise rather than an indication of recall.

[0138] Aspect 24: According to aspect 21, 22, or 23, further comprising: determining that an intent specificity score for the search query is higher than an intent specificity score threshold; based on determining that the intent specificity score for the search query is higher than the intent specificity score threshold, determining that the search query is a query-related factor; based on determining that the search query is a query-related factor, ranking a set of search results for the search query; and providing a set of search results for the search query based on the ranking.

[0139] Aspect 25: According to aspect 21, 22, 23, 24, or 25, further comprising: identifying a plurality of queries of a third query type different from the second query type and the first query type, the third query type corresponding to a higher number of user interactions compared to the second query type and a lower number of user interactions compared to the first query type; generating a second intent specificity distribution of search results for the plurality of queries of the third query type based on user interactions with the search results of the plurality of queries of the third query type; using the second intent specificity distribution to train the intent specificity machine learning model to generate an intent specificity score for search queries of the second query type; and providing an intent specificity score for the search query based on training the intent specificity machine learning model using the second intent specificity distribution.

[0140] Aspect 26: According to aspect 21, 22, 23, 24, or 25, wherein the first query type is a head query type, the second query type is a tail query type, and the third query type is a torso query type.

[0141] Aspect 27: According to aspect 21, 22, 23, 24, 25, or 26, wherein the intent specificity machine learning model includes Bidirectional Encoder Representations from Transformers (BERT).

[0142] Aspect 28: According to aspect 21, 22, 23, 24, 25, 26 or 27, wherein the first query type corresponds to a query with fewer search terms compared to a search query of the second query type.

[0143] Aspect 29: A computer system includes: a processor; and a computer storage medium storing computer-usable instructions that, when used by the processor, cause the computer system to perform operations including: identifying a first set of queries of a first query type different from a second query type, wherein the first query type corresponds to a higher number of user interactions compared to the second query type; generating an intent specificity for each query in the first set of queries based on associated search results for each query in the first set of queries and based on user interactions with the search results; using the intent specificity for each query in the first set of queries to train an intent-specific machine learning model to generate an intent specificity score for a search query of the second query type; providing a second query of the second query type to the trained intent-specific machine learning model; and using the trained intent-specific machine learning model to provide an intent specificity score for the second query.

[0144] Aspect 30: According to aspect 29, wherein the intent-specific machine learning model is trained by: generating an item list vector for each associated search result for each query in the first set of queries; using the item list vector to generate a query vector for each query in the first set of queries; and determining an average of cosine similarities between each query vector and its corresponding item list vector to generate an intent specificity for each query in the first set of queries.

[0145] Aspect 31: According to aspect 29 or 30, further comprising: determining that an intent specificity score for the second query is below an intent specificity score threshold; and based on determining that the intent specificity score for the second query is below the intent specificity score threshold, providing an indication of recall of retrieval of search results using the second query rather than an indication of precision.

[0146] Aspect 32: According to aspect 29, 30, or 31, further comprising: identifying a set of queries of a third query type that is different from the second query type and the first query type, the third query type corresponding to a higher number of user interactions compared to the second query type and a lower number of user interactions compared to the first query type; generating a query vector for each query in the set of queries of the third query type by aggregating search result embeddings of search results for each query in the set of queries of the third query type; training an intent-specific machine learning model based on the query vectors for each query in the set of queries of the third query type; and providing an intent-specific score for the second query based on training the intent-specific machine learning model using the query vectors for each query in the set of queries of the third query type.

[0147] Aspect 33: According to aspect 29, 30, 31, or 32, further comprising: determining that the intent-specific score for the second query is lower than an intent-specific score threshold; based on determining that the intent-specific score for the second query is lower than the intent-specific score threshold, determining that the second query is a query-unrelated factor; based on determining that the second query is a query-unrelated factor, ranking the search results for the second query; and providing the search results for the second query based on the ranking.

[0148] Aspect 34: According to aspect 29, 30, 31, 32, or 33, wherein the first set of queries includes at least a first query and a third query, and wherein the intent-specific machine learning model is trained based on: aggregating search result embeddings of search results for the first query to generate a first query vector; aggregating search result embeddings of search results for the third query to generate a second query vector; determining a cosine angle between the first query vector and each search result embedding of the search results for the first query to generate an intent-specificity for the first query; and determining a cosine angle between the second query vector and each search result embedding of the search results for the third query to generate an intent-specificity for the third query.

[0149] Aspect 35: One or more non-transitory computer storage media storing computer-usable instructions that, when used by a computing device, cause the computing device to perform operations including: identifying a first set of queries of a first query type different from a second query type, the first set of queries including at least a first query and a second query, the first query type corresponding to a higher number of user interactions compared to the second query type; generating a first query vector for the first query by aggregating search result embeddings of search results for the first query and based on user interactions with the search results for the first query; generating a second query vector for the second query by aggregating search result embeddings of search results for the second query and based on user interactions with the search results for the second query; training an intent-specific machine learning model based on generating the first query vector and the second query vector to generate an intent-specific score for search queries of the second query type; providing a third query of the second query type to the trained intent-specific machine learning model; and using the trained intent-specific machine learning model to provide an intent-specific score for the third query.

[0150] Aspect 36: According to aspect 35, wherein the search result embeddings of the search results for the first query include item list embeddings.

[0151] Aspect 37: According to aspect 35 or 36, further including: identifying a second set of queries of a third query type different from the second query type and the first query type, the second set of queries including at least a fourth query and a fifth query, the third query type corresponding to a higher number of user interactions compared to the second query type and a lower number of user interactions compared to the first query type; generating a third query vector for the fourth query by aggregating search result embeddings of search results for the fourth query and based on user interactions with the search results for the fourth query; generating a fourth query vector for the fifth query by aggregating search result embeddings of search results for the fifth query and based on user interactions with the search results for the fifth query; training an intent-specific machine learning model based on generating the third query vector and the fourth query vector; and using the intent-specific machine learning model trained based on generating the third query vector and the fourth query vector to provide an intent-specific score for the third query.

[0152] Aspect 38: According to aspect 35, 36 or 37, wherein the intent-specific machine learning model is trained using intent-specificity determined using the third query vector and the fourth query vector.

[0153] Aspect 39: According to aspect 35, 36, 37 or 38, it further includes: determining that the intent specificity score for the third query is higher than an intent specificity score threshold; and based on the determination that the intent specificity score for the third query is higher than the intent specificity score threshold, providing an indication that the retrieval of search results using the third query is precise, rather than providing an indication for recall.

[0154] Aspect 40: According to aspect 35, 36, 37, 38 or 39, it further includes: determining that the intent specificity score for the third query is higher than an intent specificity score threshold; based on the determination that the intent specificity score for the third query is higher than the intent specificity score threshold, determining that the third query is a query-related factor; based on the determination that the third query is a query-related factor, ranking the search results for the third query; and providing the search results for the third query based on the ranking.

Claims

1. A computer-implemented method comprising: Identify search queries performed using search engines; generating a query vector for the search query by aggregating item list vectors from search results for the search query; determining a similarity between the query vector and the item list vector; as well as The intent specificity of the search query is generated based on aggregating similarities between the query vector and the item list vector.

2. The computer-implemented method of claim 1 , wherein: The search results have a user interaction history generated by a user device that provided the search query to the search engine.

3. The computer-implemented method of claim 1 , wherein: Aggregating the item list vectors to generate the query vector includes determining an average of the item list vectors.

4. The computer-implemented method of claim 1 , wherein: The similarity between the query vector and the item list vector is determined by using the Euclidean distance.

5. The computer-implemented method of claim 1 , wherein: The similarity between the query vector and the item list vector is determined using a cosine similarity between the query vector and the item list vector.

6. The computer-implemented method of claim 5, wherein: Aggregating the similarities between the query vector and the item list vector includes determining an average of the cosine similarities.

7. The computer-implemented method of claim 1 , further comprising: identifying a second search query performed using the search engine; generating at least one query vector for the second search query by aggregating the item list vectors from the search results of the second search query; determining a similarity between the at least one query vector for the second search query and an item list vector from search results of the second search query; as well as A second intent specificity of the second search query is generated based on aggregating similarities between the at least one query vector for the second search query and item list vectors from search results of the second search query.

8. The computer-implemented method of claim 1 , further comprising: training an intent-specific machine learning model using the intent-specificity and the second intent-specificity to generate an intent-specificity score for an additional search query; providing a third query to the trained intent-specific machine learning model; and Providing an intent-specificity score for the third query using the trained intent-specific machine learning model.

9. A computer system comprising: processor; as well as A computer storage medium storing computer usable instructions that, when used by the processor, cause the computer system to perform operations including: identifying a first set of queries having a determined intent specificity; training an intent-specific machine learning model using the intent-specificity for each query in the first set of queries to generate intent-specificity scores for additional search queries; providing a second query to the trained intent-specific machine learning model; and The trained intent-specific machine learning model is used to provide an intent-specific score for the second query.

10. The computer system according to claim 9, wherein: The intent specificity of the first query set is determined by: generating an item list vector for each search result set corresponding to each query in the first query set; generating a query vector for at least one query in the first set of queries using the item list vector of search results for the at least one query in the first set of queries; as well as An average of cosine similarities between the query vector and item list vectors of search results for the at least one query in the first set of queries is determined to determine intent specificity for the at least one query in the first set of queries.

11. The computer system according to claim 9, wherein: The first set of queries each have a history of previous user interactions above a threshold with search results for each query in the first set of queries.

12. The computer system according to claim 11, wherein: The second query provided to the trained intent-specific machine learning model has a history of prior user interactions with search results for the second query that is below the threshold.

13. The computer system of claim 12, further comprising: identifying a second set of queries having a determined specificity of intent, the first set of queries having a higher number of user interactions with search results than the second set of queries, the second set of queries having a prior user interaction history above the threshold; training the intent-specific machine learning model using intent-specificity for each query in the second set of queries; as well as Based on training the intent-specificity machine learning model using the intent-specificity for each query in the second set of queries, an intent-specificity score for the second query is provided.

14. The computer system of claim 9, further comprising: determining that the intent specificity score for the second query is below an intent specificity score threshold; determining that the second query is a query irrelevant factor based on determining that the intent specificity score for the second query is below the intent specificity score threshold; ranking search results for the second query based on determining that the second query is a query irrelevant factor; as well as Search results for the second query are provided based on the ranking.

15. One or more non-transitory computer storage media storing computer usable instructions that, when used by a user device, cause the user device to perform operations comprising: Identify search queries performed using search engines; generating a query vector for the search query by aggregating item list vectors from search results for the search query; determining a similarity between the query vector and the item list vector; as well as The intent specificity of the search query is generated based on aggregating similarities between the query vector and the item list vector.

16. The one or more non-transitory computer storage media of claim 15, further comprising: identifying a second query executed using the search engine; aggregating search result embeddings of search results for the second query to generate a second query vector; A cosine angle is determined between the second query vector and each search result embedding of search results for the second query to generate intent specificity for the second query.

17. The one or more non-transitory computer storage media of claim 16, further comprising: training an intent-specific machine learning model using the intent-specificity for the search query and the intent-specificity for the second query; providing a third query to the trained intent-specific machine learning model, the third query having a lower user interaction history with search results for the third query than the first query and the second query, the intent-specific machine learning model being trained to: generate an intent-specificity score for queries having a lower user interaction history with search results than the user interaction history with the first query and the second query; as well as Providing an intent-specificity score for the third query using the trained intent-specificity machine learning model.

18. The one or more non-transitory computer storage media of claim 17, wherein: The search result embedding includes an item list embedding.

19. The one or more non-transitory computer storage media of claim 17, further comprising: determining that an intent-specificity score for the third query is above an intent-specificity score threshold; as well as Based on determining that the intent specificity score for the third query is above the intent specificity score threshold, providing an indication that retrieval of search results using the third query was precise rather than providing an indication of recall.

20. The one or more non-transitory computer storage media of claim 17, further comprising: determining that an intent-specificity score for the third query is above an intent-specificity score threshold; determining that the third query is a query-relevant factor based on determining that the intent-specificity score for the third query is above the intent-specificity score threshold; ranking search results for the third query based on determining that the third query is a query-relevant factor; as well as Search results for the third query are provided based on the ranking.