Similarity sensitive diversity

By introducing similarity sensitive diversity in search engines, measuring the distribution changes of item lists in categories, the problem of inappropriate distribution of search results in the prior art is solved, and more efficient and accurate search results return is achieved.

CN120179932APending Publication Date: 2025-06-20EBAY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411860390.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-12-17
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When existing search technologies deal with ambiguous search queries, it is difficult to return appropriately distributed search results, resulting in users requiring multiple searches to identify the required products and consuming unnecessary computing resources.

Method used

By introducing similarity sensitivity diversity in search engines, measuring the distribution variation of the item list in one or more categories, generating pairs of similarity matrices, pruning the matrix to remove category pairs below the threshold, thereby determining similarity sensitivity diversity among items.

Benefits of technology

Improves search accuracy, reduces search queries and filter selections that users need to submit, reduces computing resource consumption, such as processing power and network bandwidth, reduces storage device I/O operations, and improves the diversity and applicability of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179932A_ABST
    Figure CN120179932A_ABST
Patent Text Reader

Abstract

The similarity sensitive diversity is used to measure changes in the distribution of the list of items in one or more categories. Cosine similarities between category vectors of each category pair in the set of categories are determined, and paired similarity matrices are generated using the cosine similarities. The pairwise similarity matrix may be pruned to remove category pairs below a threshold. By utilizing the pairwise similarity matrices, a similarity-sensitive diversity between one or more of the plurality of items may be determined. In various aspects, similarity sensitive diversity may be used to generate a list of appropriately distributed related items suggesting refinement of a search query; generating a navigation module; classifying or reclassifying the plurality of items; or generating an automatic suggestion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to search engines, and more particularly to similarity-sensitive diversity. Background Art

[0002] Many product search systems allow users to submit search queries consisting of several words or terms. The search system returns a list of relevant items related to the search query based on keyword searches available within the corresponding site. If some relevant items have similar attributes but correspond to different categories, the search system may not return a properly distributed result. Summary of the Invention

[0003] At a high level, the aspects described herein relate to search engines. More specifically, the aspects described herein relate to a search engine that utilizes similarity-sensitive diversity to measure changes in the distribution of a list of items across one or more categories. The cosine similarity between the category vectors of each pair of categories in a set of categories is determined, and this cosine similarity is used to generate a pairwise similarity matrix. The pairwise similarity matrix can be pruned to remove category pairs below a threshold. By utilizing the pairwise similarity matrix, the similarity-sensitive diversity between one or more of a plurality of items can be determined. In various aspects, similarity-sensitive diversity can be used to: generate a list of relevant items with a proper distribution, suggest refinements to a search query; generate a navigation module; classify or reclassify a plurality of items; or generate autocomplete suggestions.

[0004] This Summary of the Invention is intended to introduce a selection of concepts in a simplified form that are further described in the detailed description of the present disclosure. This Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to assist in determining the scope of the claimed subject matter. Additional objects, advantages, and novel features of the present technology will be provided, and will in part become apparent to those skilled in the art upon examination of the present disclosure or upon learning by practice of the present technology. Brief Description of the Drawings

[0005] This patent or application document contains at least one color drawing. Copies of this patent or patent application publication with color drawings will be provided by the Patent and Trademark Office upon request and payment of the necessary fees. Embodiments of the present technology are described in detail below with reference to the drawings, in which:

[0006] Figure 1 is a block diagram of an example operating environment suitable for implementing aspects of the present technology;

[0007] Figure 2 is a flowchart showing an example method in accordance with aspects of the technology described herein;

[0008] Figure 3is a flowchart showing an example method in accordance with the technical aspects described herein; and

[0009] Figure 4 is an example computing device suitable for implementing the techniques in accordance with the aspects described herein. DETAILED DESCRIPTION

[0010] The subject matter of the present technical aspects is described herein in detail to meet statutory requirements. However, the present specification itself is not intended to limit the scope of this patent. On the contrary, the inventors have contemplated that the claimed subject matter may also be embodied in other ways in combination with other existing or future technologies, to include different steps or combinations of steps similar to those described in this document. Additionally, although the terms "step" and / or "block" may be used herein to denote different elements of the methods employed, such terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless and except when the order of individual steps is explicitly described.

[0011] Furthermore, unless otherwise stated, words such as "a" and "an" may also include the plural as well as the singular. Thus, for example, in the presence of one or more features, the constraint of "a feature" is satisfied. Additionally, the term "or" includes conjunctive, disjunctive, and both (thus a or b includes a or b, as well as a and b).

[0012] Although search engines are very useful tools for providing search results for received search queries, the drawbacks of existing search techniques often result in the consumption of an unnecessary amount of computing resources (e.g., I / O costs, network packet generation costs, throughput, memory consumption, etc.). When performing a search, a user typically looks for specific search results but may enter a somewhat ambiguous search query. Even if the search query is not ambiguous, the items responding to the search query may not be properly classified (the items may be classified into multiple but similar categories). For example, in the context of searching for PELOTON, a user may be looking for women's shirts of the PELOTON brand. However, the results provided to the user may include categories such as fitness bikes, women's sweatpants, sports bras, women's sports tops, women's tops, men's sports tops, etc. Many of these categories are very similar, and if these categories are combined (e.g., women's sweatpants, sports bras, women's sports tops, women's tops) and the search results are displayed according to the categories that confirm the similarity of the items, such that the user can find women's shirts of the PELOTON brand in a single search, then the search results will be more useful to the user. In this context, existing search techniques cannot provide useful or appropriately diverse search results, and useful or appropriately diverse search results subsequently require the user to submit additional search queries or multiple filters to obtain the desired search results.

[0013] This requires the user to perform multiple searches to identify products available for purchase among multiple categories. This process unnecessarily consumes various computing resources of the search system, such as processing power, network bandwidth, throughput, memory consumption, etc. In some cases, multiple attempts to identify products may not even fully meet the user's goals, so the user needs to spend more time and computing resources during the search process by repeating the process of issuing additional queries before the user finally accesses the desired content item. In some cases, the user may even abandon the search because the search engine cannot return the desired search results after multiple searches.

[0014] These drawbacks of existing search technologies have an adverse impact on computer network communication. For example, each time a query is received, the content or payload of the search query is typically supplemented with header information or other metadata, which is multiplied by all the additional queries required to obtain the specific item the user desires. Thus, repeatedly generating this metadata and sending it over the computer network incurs throughput and latency costs. In some cases, these repeated inputs (e.g., repeated clicks, selections, or queries) increase storage device I / O (e.g., excessive physical read / write head movement on non-volatile disks) because each time the user enters unnecessary information (such as entering several queries), the computing system often has to access the storage device to perform read or write operations, which is both time-consuming and error-prone, and ultimately wears out components such as read / write heads. Additionally, if multiple users repeatedly issue queries, it is costly because processing queries consumes a large amount of computing resources. For example, for some search engines, a query execution plan may need to be calculated each time a query is issued, which requires the search system to find the lowest-cost query execution plan to fully execute the query. This reduces throughput and increases network latency, and wastes valuable time.

[0015] Given these drawbacks of existing search technologies, the technical aspects described herein improve the functionality of the computer itself by providing a solution that enables a search engine to utilize similarity-sensitive diversity to measure changes in the distribution of a list of items across one or more categories. In various aspects, similarity-sensitive diversity can be used to: generate a list of relevant items with an appropriate distribution, suggest refinements to a search query; generate a navigation module; classify or reclassify multiple items; or generate autocomplete suggestions. It can be understood that better results are obtained compared to traditional search engines that require multiple search queries from the user.

[0016] The technical aspects described herein provide several improvements over existing search techniques. For example, the computational resource consumption is improved relative to the prior art. Specifically, by leveraging similarity-sensitive diversity to measure changes in the distribution of an item list across one or more categories, search accuracy is improved, thereby allowing users to access relevant search results more quickly. This eliminates (or at least reduces) duplicate user queries and filter selections, as the search results are appropriately distributed and classified. Consequently, the technical aspects described herein reduce computational resource consumption, such as processing power and network bandwidth. For example, a user query (e.g., an HTTP request) will only need to traverse the computer network once (or fewer times relative to the prior art).

[0017] In a similar manner, the technical aspects described herein improve storage device or disk I / O and query execution functionality. As described above, the insufficient search results provided by existing search techniques lead to duplicate user queries and filter selections. This results in multiple traversals of the disk I / O. In contrast, these aspects described herein reduce storage device I / O because the user provides a reduced amount of input, and thus the computing system does not have to access the storage device as frequently to perform read or write operations. For example, by leveraging similarity-sensitive diversity, a search engine can provide enhanced search results that enable users to identify and purchase items that are appropriately distributed and classified through a single search query. Consequently, there is not as much wear and tear due to the query execution functionality.

[0018] An overview of the technical aspects described herein has been briefly described. Below, an exemplary operating environment in which the technical aspects described herein can be implemented is described.

[0019] Now turning to Figure 1 , a block diagram of an operating environment 100 is provided that illustrates where these aspects of the present disclosure can be employed. It should be understood that such and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, command, and function groupings) can be used to supplement or replace the shown arrangements and elements, and some elements can be omitted altogether for clarity. Additionally, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components, and implemented in any suitable combination and location. The various functions described herein that are performed by one or more entities can be executed by hardware, firmware, and / or software. For example, some functions can be executed by a processor that executes instructions stored in a memory.

[0020] Among other components not shown, exemplary operating environment 100 includes: a network 102; a computing device 104 having a client interface component 106; a search engine 108 having a query module 110, a search module 112, and a similarity module 114; a keyword index 130; and an item database 134. It should be understood that Figure 1 the illustrated environment 100 is an example of a suitable operating environment. For example, Figure 1 each of the illustrated components can be implemented via any type of computing device (e.g., the computing device 400 described below in conjunction with Figure 4 ).

[0021] These components can communicate with each other via the network 102, which can include but is not limited to one or more local area networks (LANs) and / or wide area networks (WANs). In an exemplary embodiment, the network 102 includes the Internet and / or a cellular network, as well as any one of a variety of possible public and / or private networks. In some aspects, the network 102 can include multiple networks and one of the multiple networks, but is shown in a simpler form to avoid obscuring other aspects of the present disclosure.

[0022] It should be understood that within the scope of the present disclosure, any number of user devices, servers, and data sources can be employed within the operating environment 100. Each can include a single device or multiple devices that cooperate within a distributed environment. For example, the search engine 108 can be provided via multiple devices arranged within a distributed environment that together provide the functions described herein. Additionally, other components not shown can also be included within the distributed environment.

[0023] The computing device 104 can be a client device on the client side of the operating environment 100, while the search engine 108 can be on the server side of the operating environment 100. For example, the search engine 108 can include server-side software that is designed to work in conjunction with client-side software on the computing device 104 in order to implement any combination of the features and functions discussed in the present disclosure. This division of the operating environment 100 is provided to illustrate an example of a suitable environment, and for each embodiment, no particular combination of the search engine 108 and the computing device 104 is required to remain as separate entities. Although the operating environment 100 shows a configuration within a networked environment having separate computing devices, search engines, keyword indexes, and item databases, it should be understood that other configurations in which components are combined can be employed. For example, in some configurations, the computing device can also serve as a data source and / or can provide search capabilities.

[0024] The computing device 104 can include any type of computing device that a user can use. For example, in one aspect, the computing device 104 can be as described herein with respect toFigure 4 The type of computing device 400 described. By way of example and not limitation, a computing device may be embodied as a personal computer (PC), laptop computer, mobile device or mobile equipment, smart phone, tablet computer, smart watch, wearable computer, personal digital assistant (PDA), MP3 player, global positioning system (GPS) or device, video player, handheld communication device, gaming device or system, entertainment system, in-vehicle computer system, embedded system controller, remote control, appliance, consumer electronic device, workstation, or any combination of these described devices, or any other suitable device that can execute a search query via the client interface component 106 or can present a notification via the client interface component 106. A user may be associated with the computing device 104. The user may communicate with the search engine 108 via one or more computing devices such as the computing device 104.

[0025] At a high level, the search engine 108 receives a text-based search query (e.g., a natural language query or a structured query) or an audio query including speech or other audio input from the computing device 104 (or another computing device not depicted). In some aspects, the text-based query or audio query includes one or more keywords. The search query may include any type of input from the user for initiating a search including one or more keywords. In response to receiving the search query, the search engine 108 generates and ranks text-based results in a single search result set.

[0026] In some configurations, the search engine 108 may be embodied on one or more servers. In other configurations, the search engine 108 may be implemented at least partially or fully on a user device (e.g., Figure 4 the computing device 400 described). The search engine 108 (and its components) may be embodied as a set of compiled computer instructions or functions, program modules, computer software services, or an arrangement of processing executed on one or more computer systems.

[0027] As Figure 1As shown, the search engine 108 includes a query module 110, a search module 112, and a similarity module 114. In one aspect, the functions performed by the modules of the search engine 108 are associated with one or more applications, services, or routines. Specifically, such applications, services, or routines may run on one or more user devices (e.g., computing device 104) or servers (e.g., search engine 108), or may be distributed across one or more user devices and servers. In some aspects, the applications, services, or routines may be implemented in the cloud. Additionally, in some aspects, these modules of the search engine 108 may be distributed across a network in the cloud, which includes one or more servers and client devices (e.g., computing device 104); or may reside on a user device such as the computing device 104.

[0028] Furthermore, the modules of the search engine 108, as well as the functions and services performed by these modules, may be implemented at an appropriate level of abstraction, such as an operating system level, an application level, or a hardware level, etc. Alternatively or additionally, the functions of these modules (or the technical aspects described herein) may be performed at least in part by one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), etc. Additionally, although the functions are described herein for the specific modules shown for the search engine 108, it is contemplated that in some aspects, the function of one of the modules may be shared or distributed among the other modules.

[0029] The query module 110 receives a search query that includes one or more text-based keywords. For example, a user may enter a search query at the computing device 104 via a client interface component 106 that provides access to the search engine. As previously described, the search query may include any type of input from the user for initiating a search that includes one or more keywords.

[0030] The query module 110 may be configured to receive the search. Additionally, the query module 110 may be configured to transmit the search query to other modules of the search engine 108, such as the search module 112 or the similarity module 114. Additionally, the user query module 110 may be configured to suggest and provide refinements to the search query, generate and provide navigation modules, classify or categorize multiple items, generate and provide auto-suggestions, or provide search results to a computing device such as the computing device 104.

[0031] The query module 110 can cause one or more graphical user interface displays of various computing devices to display a search query, a suggested refinement of the search query, a navigation module, a category of items responsive to the search query, an auto-suggestion, or an item responsive to the search query. In some aspects, the query module 110 causes the client interface component 106 (through which a search query is input (e.g., by a user in a search tool on a web page)) to display the search query, the suggested refinement, the navigation module, the category of items responsive to the search query, the auto-suggestion, or the item responsive to the search query. Additionally, the query module 110 can include an application programming interface (API) that allows applications to submit search queries (and optionally other information such as user information, context information, etc.) for receipt by the search engine 108.

[0032] The search module 112 identifies search results in response to a search query processed against the item database 134, which will be described in more detail below. For example, the search module 112 can query the keyword index 130 to identify results that meet the criteria of the search query. In some aspects, the results identified in the keyword index 130 are mapped to items in the item database 134. For clarity, an item can be a list of items of a product and can include various additional information such as price, price range, quality, condition, ranking, material, brand, manufacturer, etc.

[0033] The search module 112 also ranks the search results. In some aspects, information learned from historical search sessions or user feedback is used to optimize the ranking of the search results. For example, the choices made by other users who submitted similar queries can be used to increase or decrease the ranking of individual items within the search results.

[0034] In some aspects, the feedback can be stored in a search log. The search log can be embodied on multiple databases, one or more of which include one or more hardware components that are part of the search engine 108. In some aspects, the search log is configured to store information about a user's historical search sessions, including, for example, search queries submitted by multiple users via a client interface component (e.g., the client interface component 106), search results associated with the historical search queries, a list of items of the search results, or user interactions associated with the search results (e.g., hovering, clicking, purchasing, etc.). In some embodiments, the search log stores a timestamp (e.g., day, hour, minute, second, etc.) for each user query, the search results, a list of items associated with the search results, the user's interaction with the search results, etc.

[0035] In addition, the information stored in the search log regarding historical search sessions can include other result selection information, such as subsequent filters selected in response to receiving search results and item lists. In some embodiments, the result selection information can include the time between two consecutive selections of search results, the language employed by the user, and the country in which the user may be located (e.g., based on the server used to access the search engine 108). In some implementations, other information associated with the stored historical search sessions can include the user's interaction with the rankings displayed within the item list, negative feedback displayed with the item list, and other information (e.g., whether the user clicked on or viewed a document associated with the item list). User information, including user cookies, cookie age, IP (Internet Protocol) address, user agent of the browser, etc., can also be stored in the search log. In some embodiments, the user information is recorded in the search log for the entire user session or multiple user sessions.

[0036] The keyword index 130 and the item database 134 can include data sources or data systems that are configured to make data available to any of the various components of the operating environment 100. The keyword index 130 and the item database 134 can be discrete components separate from the search engine 108, or they can be incorporated into or integrated with the search engine 108 or other components of the operating environment 100. Additionally, the item database 134 can store search results associated with search queries, and the relevant information can be indexed in the keyword index 130.

[0037] The keyword index 130 can take the form of an inverted index, but other forms are possible. The keyword index 130 stores information about items in a manner that allows the search engine 108 to efficiently identify search results for search queries. The search engine 108 can be configured to run any number of queries against the keyword index 130. According to an example embodiment, the keyword index 130 can include an inverted index that stores a mapping from text search queries to items in the item database 134.

[0038] The similarity module 114 uses similarity-sensitive diversity to measure the distribution change of objects in one or more categories. In a search context, the objects are items (e.g., items for sale), and the dimensions are category or aspect values. In some respects, the distribution can be based on demand (i.e., buyers), supply (i.e., sellers), or a combination of both. For example, a seller can list an item for sale and describe specific aspects of the item. Additionally, the seller can assign a specific category to the item. Another seller can list the same item for sale and describe different aspects of the item or assign a different category to the item. Similarly, a buyer can search for a specific item for sale and enter specific aspects of the item in the search query or search for the item in a specific category. Another buyer can search for the same item for sale and enter different aspects of the item or search for the item in a different category. Although described in terms of categories, it is envisioned that other aspects (e.g., price, color, etc.) can be used similarly, as described below.

[0039] Using similarity-sensitive diversity, the similarity module 114 is able to reconcile the subtle differences in these aspects or categories to display the appropriate distribution of items in one or more categories. For example, consider items offered in multiple colors (e.g., red, blue, yellow, green), and there are four items for sale in each color in an electronic marketplace. Next, consider items offered in multiple shades of a single color (e.g., palest blue, light blue, blue, and dark blue), and there are four items for sale in each shade in the electronic marketplace. Finally, consider items offered in a single shade of a single color (e.g., light blue), and there are sixteen such items in the electronic marketplace.

[0040] By using a conventional diversity measure such as Shannon Entropy (for context, Shannon diversity treats the values of a probability distribution as either identical or completely different), the first two examples (e.g., items offered in multiple colors and items offered in multiple shades of a single color) would each have a diversity of approximately four. However, this ignores the similarity of items offered in multiple shades of a single color, whose diversity may be more appropriately represented as being closer to one and more similar to the third example (e.g., items offered in a single shade of a single color). For clarity, although the term Shannon Entropy is used herein to describe the conventional diversity measure, it is actually typically the power of Shannon Entropy used to measure diversity (i.e., 2 to the power of Shannon Entropy).

[0041] Now consider a set of 100 shirts in four different colors. If the shirts are evenly distributed across the colors, the Shannon diversity is four. However, if two of the four colors are very close to each other, the search results will not be properly distributed. Even if the two colors are nearly identical, the distribution ratio of the four colors will be 25%. However, by using similarity-sensitive diversity, the diversity may be slightly less than three, and the distribution ratio of the three colors will be 50% / 25% / 25%, which represents a more accurate distribution among the search results.

[0042] Returning to the reference PELOTON example, an actual experiment using Shannon entropy to determine diversity yielded 14.6 categories. For example, the results provided to the user included categories such as exercise bikes, women's sweatpants, sports bras, women's athletic tops, women's tops, men's athletic tops, etc. These results were not useful to the user because the user had to search multiple categories or apply multiple filters to identify the item the user was actually searching for. However, in the actual experiment, by using similarity-sensitive diversity, the similarity module 114 determined the diversity to be 3.11. Thus, if a user is searching for a women's shirt by the PELOTON brand and the results are displayed according to categories that more appropriately identify the actual diversity between items (e.g., exercise bikes, women's clothing, men's clothing), the user is more likely to find the item in a single search.

[0043] In various aspects, similarity-sensitive diversity can be utilized to suggest refinements to search queries, generate navigation modules, classify or reclassify multiple items, or generate auto-suggestions. For example, if a user submits a search query for PELOTON but is actually looking for a women's shirt by the PELOTON brand, once the similarity module 114 determines the similarity-sensitive diversity, the query module 110 may provide a suggested refinement to the search query, such as PELOTON bikes, PELOTON women's clothing, or PELOTON men's clothing. In a similar manner, when a user is entering a search query, the query module 110 may generate auto-suggestions, such as PELOTON bikes, PELOTON women's clothing, or PELOTON men's clothing.

[0044] In some aspects, the query module 110 changes the display of the user interface to utilize the navigation module, which, when selected, displays items corresponding to, for example, PELOTON bikes, PELOTON women's clothing, or PELOTON men's clothing. For example, the query module can provide the navigation module so that the user only sees items within the category of PELOTON items that the user is actually searching for, rather than listing all items in response to a query for PELOTON. In this way, the navigation module influenced by similarity-sensitive diversity can be used to save valuable space in the user interface.

[0045] In another example, the similarity module 114 can determine similarity-sensitive diversity and suggest categories or aspects to the seller when listing items for sale. Alternatively, if an item has already been listed for sale but may contain ambiguous categories or aspects, similarity-sensitive diversity can be utilized to reclassify the item or modify the aspect. Continuing with the PELOTON example, consider a seller listing a PELOTON brand women's shirt under the women's tops category. In a search system that classifies items using Shannon diversity or without using any form of diversity metric, a user may have to conduct multiple searches and / or filters to find the PELOTON brand women's shirt. However, by leveraging similarity-sensitive diversity, the query module 110 can suggest categories to the user when listing an item or reclassify the item after the seller has listed it.

[0046] Figure 2 is a flowchart showing a method 200 for determining similarity-sensitive diversity according to the technical aspects described herein. The method 200 can be performed, for example, by Figure 1 the search engine 108. As shown at block 202, a search query is received from a user. For example, it can be received at the search engine ( Figure 1 the search engine 108) via a client interface component (such as Figure 1 the client interface component 106) from a computing device (such as Figure 1 the computing device 104).

[0047] At block 204, a plurality of items responsive to the search query are identified. Each of the plurality of items corresponds to a category. The category can be defined by the seller of each of the plurality of items. For example, a seller can list various PELOTON items under categories including fitness bikes, women's sweatpants, sports bras, women's sports tops, women's tops, men's sports tops, and the like.

[0048] At block 206, for each category, a category vector is determined. The category vector can be determined by aggregating the item vectors of the clicked or purchased items among the plurality of items. In some aspects, attributes or categories defined by the seller can be used to determine the category vector.

[0049] At block 208, the cosine similarity between the category vectors of each category pair is determined to generate a pairwise similarity matrix. Although the cosine similarity between the category vectors of each category pair is described herein and used to generate a pairwise similarity matrix, it is contemplated that, and within the scope of the present disclosure, any vector similarity function (e.g., Euclidean distance) can be similarly utilized. In some aspects, the pairwise similarity matrix is pruned to remove category pairs below a threshold (e.g., remove category pairs below, e.g., 0.88 and treat as 0). At block 210, by utilizing the pairwise similarity matrix, similarity-sensitive diversity is determined.

[0050] In some aspects, similarity-sensitive diversity can be utilized to suggest refinements to a search query. In other aspects, similarity-sensitive diversity can be utilized to generate a navigation module. In other aspects, similarity-sensitive diversity can be utilized to classify or reclassify a plurality of items. In other aspects, similarity-sensitive diversity can be utilized to generate autocomplete suggestions.

[0051] Figure 3 is a flowchart showing a method 300 for determining similarity-sensitive diversity in accordance with technical aspects described herein. Method 300 can be performed, for example, by Figure 1 search engine 108. As shown at block 302, for each category in a set of categories, a category vector is determined.

[0052] At block 304, the cosine similarity between the category vectors of each category pair is determined to generate a pairwise similarity matrix. At block 306, the pairwise similarity matrix is pruned to remove category pairs below a threshold. At block 308, the pairwise similarity matrix is utilized to determine similarity-sensitive diversity among one or more of a plurality of items. Each of the one or more items corresponds to a category in the set of categories. For example, the pairwise similarity matrix Z is used as an input that describes the similarity between each pair of values. If Z is the identity matrix, then each pair of distinct values is considered completely unrelated. Non-diagonal non-zero values indicate the similarity accounted for by similarity-sensitive diversity. Z can be used to adjust the definition of information content to make it similarity-sensitive: . Then, similarity-sensitive diversity can be determined by: .

[0053] Referring to Figure 4 , computing device 400 includes a bus 410 that directly or indirectly couples the following devices: a memory 412, one or more processors 414, one or more presentation components 416, one or more input / output (I / O) ports 418, one or more I / O components 420, and an illustrative power supply 422. Bus 410 represents one or more buses (e.g., an address bus, a data bus, or a combination thereof). Although for clarityFigure 4 The various boxes are shown by lines, but in reality, these boxes represent logical components, not necessarily physical components. For example, a rendering component such as a display device can be considered an I / O component. Additionally, the processor has a memory. The inventors of the present invention recognize this as the nature of the art and reiterate Figure 4 the figures only illustrate exemplary computing devices that can be used in conjunction with one or more aspects of the present technology. No distinction is made among these categories such as "workstation", "server", "laptop computer", "handheld device", etc., because all of these categories are within Figure 4 the scope and are considered with reference to "computing device".

[0054] Computing device 400 generally includes various computer-readable media. Computer-readable media can be any available media that can be accessed by computing device 400 and includes volatile and non-volatile media, removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media.

[0055] Computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media includes but is not limited to: RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage devices, magnetic tape cartridges, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other media that can be used to store the desired information and can be accessed by computing device 400. Computer storage media does not include signals per se.

[0056] Communication media typically embodies computer-readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transmission mechanism and includes any information delivery media. The term "modulated data signal" refers to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example and not limitation, communication media includes wired media such as a wired network or direct wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Combinations of any of the above should also be included within the scope of computer-readable media.

[0057] The memory 412 includes computer storage media in the form of volatile and / or non-volatile memory. The memory can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid state memory, hard disk drives, optical disk drives, and the like. The computing device 400 includes one or more processors 414 that read data from various entities such as the memory 412 or the I / O component 420. The presentation component 416 presents data indications to a user or other device. Exemplary presentation components include display devices, speakers, printing components, vibration components, and the like.

[0058] The I / O port 418 allows the computing device 400 to be logically coupled to other devices including the I / O component 420, some of which may be built-in. Illustrative components include microphones, joysticks, gamepads, satellite dishes, scanners, printers, wireless devices, and the like.

[0059] The I / O component 420 may provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by a user. In some instances, the inputs may be sent to an appropriate network element for further processing. The NUI may implement any combination of the following: speech recognition, touch and stylus recognition, facial recognition, biometric recognition, gesture recognition both on and near the screen, air gestures, head and eye tracking, and touch recognition associated with a display on the computing device 400. The computing device 400 may be equipped with a depth camera such as a stereoscopic camera system, an infrared camera system, an RGB camera system, and combinations thereof for gesture detection and recognition. Additionally, the computing device 400 may be equipped with an accelerometer or gyroscope capable of detecting motion. The output of the accelerometer or gyroscope may be provided to the display of the computing device 400 to present an immersive augmented reality or virtual reality.

[0060] Some aspects of computing device 400 may include one or more radios 424 (or similar wireless communication components). The radio 424 transmits and receives radio or wireless communications. The computing device 400 may be a wireless terminal adapted to receive communications and media via various wireless networks. The computing device 400 may communicate with other devices via wireless protocols such as Code Division Multiple Access (“CDMA”), Global System for Mobile Communications (“GSM”), or Time Division Multiple Access (“TDMA”). The radio communications may be short-range connections, long-range connections, or a combination of both short-range radio telecommunications connections and long-range radio telecommunications connections. When we refer to “short” and “long” types of connections, we do not mean the spatial relationship between two devices. Instead, we generally refer to short-range and long-range as different categories or types of connections (i.e., primary connections and secondary connections). By way of example and not limitation, a short-range connection may include a Wi-Fi® connection to a device that provides access to a wireless communication network (e.g., a mobile hotspot), such as a WLAN connection using the 802.11 protocol; a Bluetooth connection to another computing device is a second example of a short-range connection or a near-field communication connection. By way of example and not limitation, long-range connections may include connections using one or more of the CDMA, GPRS, GSM, TDMA, and 802.16 protocols.

Claims

1. A method for determining similarity-sensitive diversity, the method comprising: receiving a search query from a user; identifying a plurality of items responsive to the search query, the items corresponding to categories; For each category, determine a category vector; determining the cosine similarity between the category vectors for each category pair to generate a pairwise similarity matrix; The similarity-sensitive diversity is determined by utilizing the pairwise similarity matrix.

2. The method according to claim 1, further comprising: The pairwise similarity matrix is ​​pruned to remove category pairs below a threshold.

3. The method according to claim 1, wherein: The category vector is determined by aggregating item vectors of clicked or purchased items among the plurality of items.

4. The method according to claim 1, further comprising: The similarity-sensitive diversity is utilized to suggest refinements of the search query.

5. The method according to claim 1, further comprising: The similarity-sensitive diversity is used to generate a navigation module.

6. The method according to claim 1, further comprising: The plurality of items are classified using the similarity-sensitive diversity.

7. The method according to claim 1, further comprising: The similarity-sensitive diversity is exploited to generate automatic suggestions.

8. One or more non-transitory computer storage media storing computer readable instructions that, when executed by a processor, cause the processor to perform operations comprising: For each category in the category set, determining a category vector; determining the cosine similarity between the category vectors for each category pair to generate a pairwise similarity matrix; pruning the pairwise similarity matrix to remove category pairs below a threshold; A similarity-sensitive diversity between one or more items in a plurality of items is determined by utilizing the pairwise similarity matrix, each item in the one or more items corresponding to a category in the set of categories.

9. The medium of claim 8, further comprising: The pairwise similarity matrix is ​​pruned to remove category pairs below a threshold.

10. The medium according to claim 8, wherein The category vector is determined by aggregating item vectors of clicked or purchased items among the plurality of items.

11. The medium of claim 8, further comprising: The similarity-sensitive diversity is utilized to suggest refinements of the search query.

12. The medium of claim 8, further comprising: The similarity-sensitive diversity is used to generate a navigation module.

13. The medium of claim 8, further comprising: The plurality of items are reclassified using the similarity-sensitive diversity.

14. The medium of claim 8, further comprising: The similarity-sensitive diversity is exploited to generate automatic suggestions.

15. A system for determining similarity-sensitive diversity, the system comprising: at least one processor; as well as One or more computer storage media storing computer readable instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising: receiving a search query from a user; identifying a plurality of items responsive to the search query, the items corresponding to categories; For each category, determine a category vector; determining the cosine similarity between the category vectors for each category pair to generate a pairwise similarity matrix; The similarity-sensitive diversity is determined by utilizing the pairwise similarity matrix.

16. The system of claim 15, further comprising: The pairwise similarity matrix is ​​pruned to remove category pairs below a threshold.

17. The system of claim 15, further comprising: The similarity-sensitive diversity is utilized to suggest refinements of the search query.

18. The system of claim 15, further comprising: The similarity-sensitive diversity is used to generate a navigation module.

19. The system of claim 15, further comprising: Also included is classifying the plurality of items using the similarity-sensitive diversity.

20. The system of claim 15, further comprising: The similarity-sensitive diversity is exploited to generate automatic suggestions.