User query search result refinement using search logs

WO2026178728A1PCT designated stage Publication Date: 2026-09-03PAYPAL INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/079167
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2026-09-03

Smart Images

  • Figure CN2025079167_03092026_PF_FP_ABST
    Figure CN2025079167_03092026_PF_FP_ABST
Patent Text Reader

Abstract

System and methods for providing a set of results in response to a user query can include obtaining text inputs for user queries, retrieving a first log data of the user queries, determining embedding vectors representative of the user queries and first log data, partitioning the embedding vectors into a plurality of clusters, mapping each cluster to an electronic document, generating a first index representative of a frequency each electronic document is queried by the users and a second index representative of the mapped electronic documents, and storing the first index and the second index in the system to fulfill subsequent user queries. The system may, in response to a subsequent user query from a user, retrieve a second log data of the user and apply the second log data to the first and second index to provide a set of ranked results in response to the subsequent user query.
Need to check novelty before this filing date? Find Prior Art

Description

USER QUERY SEARCH RESULT REFINEMENT USING SEARCH LOGSFIELD

[0001] The present disclosure relates to the field of user search query fulfillment. More particularly, to search result refinement using user query and search logs.BACKGROUND

[0002] In an electronic transaction system of an entity, users can have user accounts in the electronic transaction system or computing network associated with the electronic transaction system so that they can perform electronic transactions on the system. A user’s account can be linked to one or more external accounts on a corresponding external computing network associated with another entity. These external accounts can be, for example, the user’s account or an account associated with another party on the external computing network. Performing transactions in such electronic transaction systems can oftentimes include performing one or more electronic transactions with these external accounts.

[0003] To link the user’s account in the electronic transaction system with an external account, a user interface is typically provided to the user such as, for example, on a user computing device. In response to a user’s input at the user interface, the user interface can typically identify a set of entity names that match the user’s input and provide the set of entity names, or a portion thereof, to the user. Accordingly, to proceed with linking the external account with the user’s account in the electronic transaction system, the user can provide a further input at the user interface by selecting one of the set of entity names that is associated with the user’s external account.BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Some embodiments of the disclosure are herein described, by way of example only, with reference to the accompanying drawings. With specific reference now to the drawings in detail, it is stressed that the embodiments shown are by way of example and for purposes of illustrative discussion of embodiments of the disclosure. In this regard, the description taken with the drawings makes apparent to those skilled in the art how embodiments of the disclosure may be practiced.

[0005] FIG. 1 is a block diagram of an example system including a search system, according to some embodiments.

[0006] FIG. 2 is a flow diagram of an example method for providing a set of results in response to a user query, according to some embodiments.

[0007] FIG. 3 is a block diagram of an example system for performing the method of FIG. 2, according to some embodiments.

[0008] FIG. 4 is a flow diagram of an example method for providing a set of results in response to a user query, according to some embodiments.

[0009] FIG. 5 is a flow diagram of an example method for refining clusters representative of the user queries, according to some embodiments.

[0010] FIG. 6 is a flow diagram of an example method for mapping the clusters of user queries to an index, according to some embodiments.

[0011] FIG. 7 is a block diagram of an example system for mapping the clusters of user queries to an index, according to some embodiments.

[0012] FIG. 8 is a flow diagram of an example method for providing a set of results in response to a user query, according to some embodiments.

[0013] FIG. 9 is a block diagram of an example system for providing the set of results in response to the user query, according to some embodiments.

[0014] FIG. 10 is a block diagram of an example computing network device, according to some embodiments.

[0015] FIG. 11 is a block diagram of the example system of FIG. 1, according to some embodiments.DETAILED DESCRIPTION

[0016] Various embodiments of the present disclosure relate to systems and methods for performing a processing workflow in a search system to provide an electronic document from a plurality of electronic documents in response to a user query. The processing workflow performed by the search system can include, for example, retrieving log data for users associated with the user queries, analyzing the log data using one or more models, generating a first index representative of a frequency each of the plurality of electronic documents is queried by the plurality of users and a second index representative of a mapping of the user queries to the plurality of electronic documents, and storing the first index and the second index to fulfill subsequent user queries. Accordingly, in response to a subsequent user query from a user, the search system can apply a log data associated with the user of the subsequent user query to provide a subset of the plurality of electronic documents in response to the subsequent user query, the subset of the plurality of electronic documents being determined based on the first index and the second index.

[0017] According to some embodiments, the search system can be utilized in the context of an electronic transaction system of an entity. The search system can be utilized to enable users to link their account in the electronic transaction system with an external account in an external computing network or system of an external entity. The external account can be, for example and without limitation, an account of the user on the external computing network. To facilitate a user performing a search using the search system, a user interface can be provided to the user at the user’s computing device. At the user interface, the system can provide a predicted set of entity names based on the first index and the second index in response to a user query. The user can then provide feedback data in the form of user behavior data (i.e., user click data) at the user interface by selecting one of the set of entity names that is a match to their query to then proceed with linking the external account with the user account.

[0018] According to some embodiments, the search system can be utilized in the context of a digital wallet. In the system associated with the entity, a digital wallet can be associated with each user. The digital wallet can, for example, be associated with one or more goods or services that the user associated with the digital wallet subscribes to from the entity. In the user’s digital wallet, the search system can be utilized by the process flows for each of the one or more goods and services associated with the entity to link the user’s account with an external account.

[0019] The embodiments of the present disclosure overcome limitations of known search systems including failure to identify an entity name that matches the user query (i.e., drop-off) due to a vocabulary mismatch between the text input from the user query and the text data of the electronic documents that are representative of the entities. This drop-off is particularly evident when customers search for an entity using common aliases. The drop-off can impact, for example, up to approximately 50%of users. Accordingly, the failure to provide the correct result in response to the user query results in a negative user experience and the drop-off also delays users from being able to perform electronic transactions in the system.

[0020] The embodiments of the present disclosure provide a search system that utilizes one or more models to provide a set of results in response to a user query that provides one or more improvements over known search methods including providing, for example, low cost of implementation, low latency, multiple platform compatibility, and alias / abbreviation / typo tolerance, among other benefits. The system provides these and other improvements by processing user queries to generate the first index and the second index and then utilizing the first index and the second index to fulfill subsequent user queries. The search system also provides improvements over known search methods by providing a continuous improvement loop that captures relevant queries and reflects search trends by analyzing search logs and regularly updating the first index and the second index, as will be further described herein.

[0021] Although the embodiments of the present disclosure describe the electronic documents in the context of entities. It is to be appreciated, however, that the electronic documents are not intended to be limited to representing entities and can be representative of, for example and without limitation, entity data, user data, user account data, policy data, product data, catalog data, etc. For example, the electronic documents can be representative of entity data for financial institutions. For example, the electronic documents can be representative of entity data for users associated with the user computing devices 108 performing transactions using transaction processing system 106. For example, the electronic documents can be representative of entity data for users associated with user computing devices 110 performing transactions using transaction processing system 106. For example, the electronic documents can be representative of entity data for sanctioned user accounts. For example, the electronic documents can be representative of catalog data for inventory items being transacted using transaction processing system 106.

[0022] Among those benefits and improvements that have been disclosed, other objects and advantages of this disclosure will become apparent from the following description taken in conjunction with the accompanying figures. Detailed embodiments of the present disclosure are disclosed herein; however, it is to be understood that the disclosed embodiments are merely illustrative of the disclosure that may be embodied in various forms. In addition, each of the examples given regarding the various embodiments of the disclosure which are intended to be illustrative, and not restrictive.

[0023] FIG. 1 is a block diagram of an example system 100 including a search system 102, according to some embodiments.

[0024] The system 100 can include a search system 102, a source of data store 104, a transaction processing system 106, a plurality of user computing devices 108 (two such user devices 108a, 108b are shown) , and one or more other user computing devices 110.

[0025] The search system 102 is configured to receive a user query such as, for example, from a user computing device 108, and provides a set of results in response to the user query.

[0026] The data store 104 can include a transaction data, a log data, an electronic documents data, a user behavioral data, an index data, other data, or any combination thereof. The log data can include, but is not limited to, account data, application data, location data, language preference data, other data, or any combination thereof, associated with the users of transaction processing system 106 or search system 102. The data store 104 can further include historical data. The historical data can include a historical transaction data, a historical log data, a historical electronic documents data, a historical user behavioral data, a historical index data, other data, or any combination thereof.

[0027] The user computing devices 108 and the user computing devices 110 can be in electronic communication with transaction processing system 106 and with each other over a network 112. The user computing devices 108 can be associated with, for example, users performing electronic transactions using transaction processing system 106. The user computing devices 110 can be associated with, for example, merchants performing electronic transactions with users associated with the user computing devices 108. In some embodiments, the user computing devices 110 can be associated with, for example, users associated with an entity of the system 100. The search system 102, the data store 104, and the transaction processing system 106 can also be in electronic communication with each other via the network 112 and / or via another network.

[0028] The present disclosure refers to users queries, log data, electronic documents, users, entities, electronic transactions, and other electronic activity. Such users can be users having user accounts common to a particular service provider, a particular network, a particular electronic activity processor, etc. For example, the user accounts can be accounts with the transaction processing system 106, and the users may be legitimate users associated with those accounts. The electronic transactions and other activity may be transactions processed by, or other activity in or through, the transaction processing system 106, and / or transactions and activity outside of the transaction processing system 106. Although this disclosure refers to transactions as context for the novel methods and systems, it should be understood that such methods and systems may be applied to or in the context of a wide variety of computing actions, some of which may not be considered transactions. For example, where past transactions are considered herein, past computing actions may more broadly be considered. Similarly, where present transactions are responded to herein, present computing actions may more broadly be responded to. Similarly, where past user interactions are considered herein, past computing actions are more broadly considered.

[0029] Users may initiate transactions, review transactions, complete transactions, etc. on user computing devices 108 through the transaction processing system 106. Accordingly, the transaction processing system 106 can receive, from user computing devices 108, instructions to initiate a transaction, an instruction to accept or complete a transaction, an instruction to review one or more transactions, an instruction to retract a transaction, etc., and can respond by performing or facilitating the requested user action. Accordingly, user activity as discussed herein can include transactions instructed through the transaction processing system 106, in some embodiments, and / or user activity on one or more platforms, networks, etc. Such transactions can include linking a user account to an external account with an external computing network, or a sub-process thereof. Such transactions can further include, for example, a computing transaction such as a file creation, a revision to a file, an electronic communication, a financial transaction (or component thereof) , a real-estate transaction (or component thereof) , a service request, or any other electronic transaction. Additionally, or alternatively, user activity according to the present disclosure may be or may include an event associated with a user, such as a user navigation to a webpage, a user search request, a user selecting a search result, etc. For example, a user may select a particular search result from a set of results provided to the user computing device 108 through an electronic user interface.

[0030] The transaction processing system 106 may be associated with a particular electronic user interface and / or platform through which a user performs electronic transactions. The electronic user interface may be embodied in a website, mobile application, etc. According, the transaction processing system 106 may be associated with or wholly or partially embodied in one or more servers, which server (s) may host the interface, and through which the user computing devices 108 may access the user interface.

[0031] The user computing devices 108 may be respectively associated with different user accounts. That is, user computing device 108a may be associated with a first user account, and user computing device 108b may be associated with a second user account. Where user computing devices are discussed herein, it may be assumed that different devices are associated with different user accounts for convenience of description, though of course a single user account may be accessed from multiple devices in practical use.

[0032] According to some embodiments, a user performing a transaction on transaction processing system 106 can include utilizing search system 102 to fulfill the transaction. For example, the transaction being performed by the user on transaction processing system 106 can be a processing workflow for linking the user’s account with an external account with an external computing system or network, and the search system 102 can be utilized to provide a set of results in response to a user query at the user interface.

[0033] The search system 102 includes a processor 116 and a non-transitory, computer-readable memory 118 that contains instructions that, when executed by the processor 116, cause the search system 102 to perform one or more of the steps, processes, methods, operations, etc. described herein with respect to the search system 102. The search system 102 includes one or more functional modules embodied in the memory. According to some embodiments, the functional modules of the search system 102 can include an interface module 120, a data module 122, an embedding module 124, a cluster module 126, a mapping module 128, an index module 130, and a machine learning module 132.

[0034] According to some embodiments, the search system 102 can operate in an online or offline mode. In some embodiments, the search system 102 can perform the steps for generating the first index and the second index in an offline mode. In some embodiments, the search system 102 can perform the steps of fulfilling a user query in an online mode. In some embodiments, the search system 102 can fulfill subsequent user queries in an online mode.

[0035] According to some embodiments, the search system 102 can utilize the one or more functional modules 120, 122, 124, 126, 128, 130 to perform one or more steps for processing a user query in an online mode. In some embodiments, in the online mode, the search system 102 can obtain a user query, combine the user query with log data (e.g., second log data) to retrieve an appropriate index, provide a ranked set of results at the user interface for selection, receive a user selection of one of the set of results at the user interface, or any combination thereof. The search system 102 can provide the ranked set of results using a machine learning model. In some embodiments, the machine learning model can include a search algorithm. In some embodiments, the machine learning model can include a fuzzy search algorithm, a prefix search algorithm, a precise search algorithm, or any combination thereof. The fuzzy search algorithm identifies relevant results between the user query and the ranked second index that may include spelling errors, variations, or inconsistencies. The prefix search algorithm, in some embodiments, identifies relevant results by matching the beginning (i.e., prefix) of the user query with the aliases of the ranked second index. The prefix search algorithm, in some embodiments, identifies relevant results by matching the beginning (i.e., prefix) of the aliases of the ranked second index with the user query. The precise search algorithm, in some embodiments, identifies those results from the ranked second index that most closely matches the user query. In some embodiments, the precise search algorithm can return results that contain the exact keywords in the same order as the user query. In some embodiments, the precise search algorithm can return results that most closely matches the keywords as defined by the user query.

[0036] The search system 102 can include an interface module 120. The interface module 120 can provide an electronic user interface to a user computing device such as, for example, user computing device 108. The user interface can be displayed at the user computing device 108. The user can provide inputs at the user interface, and the search system 102 can receive the user’s inputs from the user interface. The user inputs can be, for example, a user query for an entity. The user inputs can be, for example, a user behavioral data including a user click data representative of the user’s selection one of a set of results provided to the user interface by the user in response to the user query. The interface module 120 can provide, for example, a user interface to the user computing device 108. The interface module 120 can provide, for example, the user interface to the user computing device 108 via transaction processing system 106. Based on the interface module 120 receiving the input at the user interface, the one or more other functional modules of search system 102 can perform one or more operations in accordance with the present disclosure. In some embodiments, the interface module 120 can obtain a text input for user queries from a plurality of users of transaction processing system 106. Each of the user queries can correspond to a request for one of a plurality of electronic documents.

[0037] The search system 102 can include a data module 122. The data module 122 retrieves data from a data store based on a user query. For example, the data module 122 can retrieve, in response to a user query from a user, log data associated with the user. In some embodiments, the data module 122 can retrieve data from the data store based on operations by the one or more functional modules of search system 102. In some embodiments, the data module 122 can retrieve log data from a data store. In some embodiments, the data module 122 can retrieve electronic documents from the data store.

[0038] The data module 122 can store data from the search operations of search system 102. The search system 102 can store the data generated from the user queries. In some embodiments, the data module 122 can store the indexes generated from the user queries to be retrieved to fulfill subsequent user queries corresponding to requests for return of at least one of the plurality of electronic documents from users. In some embodiments, the data module 122 can store a first index generated by the index module 130. In some embodiments, the data module 122 can store a second index generated by the index module 130. In some embodiments, the data module 122 can store a subset of the ranked electronic documents generated to fulfill a subsequent user query as in index.

[0039] The data module 122 can retrieve log data from the data store based on the user query. In some embodiments, the data module 122 can retrieve a first log data based on the user queries. In some embodiments, the first log data includes user behavior data associated with the users of the user queries. In some embodiments, the data module 122 can retrieve a second log data including user behavior data associated with a given user. In some embodiments, the data module 122 retrieves data representative of the electronic documents from the data store. In some embodiments, the data module 122 retrieves at least one of the stored indexes from the data store. In some embodiments, the data module 122 can retrieve the data from computer-readable memory 118. In some embodiments, the data module 122 can retrieve the data from data store 104. In some embodiments, the data module 122 can retrieve the data from another data store associated with search system 102. In some embodiments, the data module 122 can retrieve the data from a data store external to search system 102 and / or external to system 100.

[0040] The log data can include, but is not limited to, an account data, an application data, a location data, a language preference data, a historical user data, other user data, or any combination thereof. In some embodiments, the account data can include at least one of the application data, location data, language preference data, historical user data, other user data, or any combination thereof.

[0041] The account data can include personal data of the user. The application data can include applications (e.g., software applications) stored in the user computing device 108. For example, an application stored in the user computing device 108 can be associated with an external computing network or external entity. For example, an application stored in the user computing device 108 can be associated with a financial institution.

[0042] The location data can include a geographic location of the user or user computing device 108. For example, in some embodiments, the location data can include a country where the user is located. For example, in some embodiments, the location data can include a state or province. For example, in some embodiments, the location data can include a city. For example, in some embodiments, the location data can include an area within a certain defined radius.

[0043] The language preference data can include one or more languages. The user can, for example, designate a first language as a primary language and a second language as a secondary language, and so on, and the log data of the user can include one or more of the languages designated by the user.

[0044] The historical user data can include, but is not limited to, historical transaction data, historical user behavioral data, historical location data, historical account data, historical external account data, other like data, or any combination thereof. For example, in some embodiments, the historical user data can include text inputs from previous user queries from the user to link the user’s account to another account. For example, in some embodiments, the historical user data can include the user’s selection from a set of results provided to the user via the user interface in response to a user query. For example, in some embodiments, the historical user data can include text data (e.g., entity name) representative of external computing networks associated with external entities that the user performed electronic transactions with using transaction processing system 106.

[0045] The data module 122 can retrieve electronic documents from the data store. In some embodiments, the data module 122 can retrieve a plurality of electronic documents from a third index. In some embodiments, the third index can be a global index of electronic documents stored in the data store. In some embodiments, the data module 122 can retrieve the electronic documents from the third index based on a mapping of the user queries to the electronic documents in the third index by the mapping module 128.

[0046] The electronic documents can include text data including label data, a location, a unique identifier, alias data, or any combination thereof. In some embodiments, the label data can be a name of an entity associated with the electronic document. In some embodiments, the location data can be a location of an entity associated with the electronic document. In some embodiments, the location data can include each of the locations of an entity associated with an electronic document. For example, the entity can be a financial institution with multiple brick and mortar locations performing electronic transactions in the external computing network of the entity. For example, the entity can be a merchant with multiple brick and mortar locations performing electronic transactions in the external computing network of the entity. In some embodiments, the unique identifier can be associated with the entity. In some embodiments, the unique identifier can be associated with one of the multiple brick and mortar locations. For example, in some embodiments, the unique identifier can be a routing number. For example, in some embodiments, the unique identifier can be a SWIFT code. For example, in some embodiments, the unique identifier can be an inventory number. For example, in some embodiments, the unique identifier can be a serial number of an inventory item. In some embodiments, the alias data can include one or more aliases for the name of the electronic document, each alias being from the text input data of historical user queries having user click data linked to the electronic document. In some embodiments, the alias data can include the most common aliases for the name of the electronic document.

[0047] The search system 102 can include the embedding module 124. The embedding module 124 transforms data into embedding vectors using a machine learning model. In some embodiments, the embedding module 124 transforms data including the user query and the log data into embeddings vectors. In some embodiments, the embedding module 124 transforms data of the electronic documents into embedding vectors using a machine learning model. In some embodiments, the embedding module 124 determines embedding vectors representative of the user queries based on the text input and the first log data using the machine learning model. In some embodiments, the embedding module 124 determines embedding vectors representative of the user queries based on the text input and the second log data using the machine learning model.

[0048] The embedding module 124 transforms the text input data from the user queries into the embedding vectors using the machine learning model. In some embodiments, the text input data of a user query can be an entity name. In some embodiments, the text input data of a user query can be representative of an entity name. For example, in some embodiments, the text input data can be an alias of an entity’s name. In some embodiments, there can be a vocabulary mismatch between the text input data from the user query and the text data of the electronic document representative of the entity name stored in the search system 102.

[0049] The embedding module 124 transforms the log data into the embeddings vectors using the machine learning model. In some embodiments, the embedding module 124 transforms the text data representative of the log data into the embeddings vectors. The embedding module 124 can associate the embeddings vectors representative of the log data with the user queries to enable the search system 102 to learn the context of the user queries. In some embodiments, for example, the embedding module 124 transforms text data representative of the account data into embeddings vectors that are associated with the user queries. In some embodiments, for example, the embedding module 124 transforms text data representative of the application data into embeddings vectors that are associated with the user queries. In some embodiments, for example, the embedding module 124 transforms text data representative of the location data into embeddings vectors that are associated with the user queries. In some embodiments, for example, the embedding module 124 transforms text data representative of the historical user data into embeddings vectors that are associated with the user queries.

[0050] The embedding module 124 can provide the embeddings vectors in a multi-dimensional embedding space. The embedding module 124 thereby utilizes the machine learning model to convert the user queries into a dense embeddings vectors representation that captures the semantic relationships and contextual meaning of the user queries. The search system 102 can utilize the embeddings vectors provided by the embedding module 124 to understand the user intent by mapping the embeddings vectors in the embedding space.

[0051] According to some embodiments, the machine learning model includes an embedding model. In some embodiments, the machine learning model is an embedding model that utilizes a transformer algorithm to generate the embeddings based on the relationship between words, or parts of words, in the text input. The transformer model can be trained on a large text corpora and tuned to perform specific tasks in accordance with the present disclosure.

[0052] According to some embodiments, the machine learning model includes a word embedding model to understand a semantic relationship of the user queries. For example, in some embodiments, the machine learning model is a Word2Vec model. For example, in some embodiments, the machine learning model is a Global Vectors for Word Representations (GloVe) model. In some embodiments, the machine learning model utilizes a continuous bag-of-words algorithm. In some embodiments, the machine learning model utilizes a skip-gram algorithm. It is to be appreciated that the machine learning model is not intended to be limited to these word embedding models and can include any of a plurality of other word embedding models to understand a semantic relationship of the user queries.

[0053] According to some embodiments, the machine learning model includes a contextual word embeddings model to understand a context of the user queries. For example, in some embodiments, the machine learning model includes a Bidirectional Encoder Representations from Transformers (BERT) model. For example, in some embodiments, the machine learning model includes a Embeddings from Language Models (ELMo) model. It is to be appreciated that the machine learning model is not intended to be limited to these contextual word embedding models and can include any of a plurality of other contextual word embedding models to understand the context of the user queries.

[0054] According to some embodiments, the machine learning model can include a word embeddings model and a contextual word embeddings model. It is to be appreciated that the machine learning model is not limited to word embedding models and can include other models including, but not limited to, document embedding models, paragraph vector models, image embedding models, graph embedding models, and the like.

[0055] The search system 102 can include the cluster module 126. The cluster module 126 partitions the embeddings vectors in the embeddings space into clusters based on the vectors values using the machine learning model. The cluster module 126 groups the user queries into respective clusters according to the embeddings vectors, the embeddings vectors being representative of the semantic characteristics of the user queries. The cluster module 126 and the search system 102 can identify patterns in user behavior according to the clustering of the embeddings vectors. In some embodiments, the cluster module 126 partitions the embedding vectors representative of the user queries into a respective one of a plurality of clusters according to characteristics of each of the embedding vectors. In some embodiments, each cluster can include one or more of the user queries.

[0056] According to some embodiments, the machine learning model utilized by the cluster module 126 can include a clustering model to group the user queries into corresponding clusters according to their embeddings vectors values. The cluster module 126 can group the user queries that have the same or similar semantic characteristics as the other user queries in a cluster using the machine learning model. In this regard, the cluster module 126 groups the user queries identified as being for a common result (e.g., entity name) together using the machine learning model. In some embodiments, the mapping module 128 partitions the embedding vectors into a respective one of a plurality of clusters according to characteristics of each of the embedding vectors. In some embodiments, each cluster can include one or more of the user queries.

[0057] In some embodiments, the clustering model utilizes a hierarchical clustering algorithm, density-based clustering algorithm, distribution-based clustering algorithm, K-means clustering algorithm, other clustering algorithms, or any combination thereof. For example, the cluster module 126 can cluster the data points representative of the user queries using a K-means clustering algorithm. For example, the cluster module 126 can cluster the data points representative of the user queries using a Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm. It is to be appreciated that the machine learning model is not intended to be limited to these clustering models and can include any of a plurality of other clustering models to group the user queries into clusters according to the embeddings vectors.

[0058] According to some embodiments, the machine learning model can include a hierarchical clustering algorithm. The hierarchical clustering algorithm generates a hierarchy of clusters using the embeddings vectors. The clusters can be represented as a tree, with the root containing all the samples and the leaves containing clusters with one or more of the samples.

[0059] According to some embodiments, the machine learning model can include a density-based clustering algorithm. The density-based clustering algorithm identifies clusters based on a density of the embedding vectors in the embeddings space. The density-based clustering algorithm can use a distance metric and a minimum number of data points to group the data points together.

[0060] According to some embodiments, the machine learning model can include a distribution-based clustering algorithm. The distribution-based clustering algorithm calculates a probability from a distribution and groups the data points into clusters based on a likelihood the data point belongs to a specific probability.

[0061] According to some embodiments, the machine learning model can include a K-means clustering algorithm. The K-means clustering algorithm uses the distance between points as a criterion for cluster formation and groups data objects into a predetermined number of clusters according to these distances.

[0062] According to some embodiments, the cluster module 126 can include a filter module to filter out data points representative of those user queries that are outliers or that include random click data according to the embeddings vector values. By filtering out these data points, the cluster module 126 improves the overall data quality so that relevant and significant user interactions from the user queries are retained for further analysis by the search system 102. In some embodiments, the cluster module 126 can identify these user queries because they are not grouped into a cluster. In some embodiments, the cluster module 126 can identify these user queries because the distance of these data points from a cluster or from a central node of the cluster exceeds a threshold limit. For example, in some embodiments, the cluster module 126 can determine that a data point representative of a user query is an outlier due to a difference between the embeddings vector values of the user query and the embedding vector values of the other user queries in a nearest cluster exceeding a defined threshold. For example, in some embodiments, the cluster module 126 can determine that a data point representative of a user query is a random click due to the embeddings vector values representative of the user click data for the user query having a different value from the embedding vector values representative of the user click data for the other user queries.

[0063] The search system 102 includes mapping module 128. For user queries that include user click data, the mapping module 128 associates (i.e., links) each of these user queries to an electronic document of a plurality of electronic documents based on the corresponding user click data. In some embodiments, the mapping module 128 maps each of the plurality of clusters to a respective one of the electronic documents according to the partitioned embedding vectors associated with each cluster. In some embodiments, the mapping module 128 associates each of the user queries remaining from the filtering step to an electronic document of a plurality of electronic documents based on the corresponding user click data. In some embodiments, the mapping module 128 associates the remaining user queries to an electronic document of the plurality of electronic documents based on embedding vector values representative of the corresponding user click data. In this regard, the user click data for a user query can be representative of the user’s selection at the user interface of one of a set of results provided to the user computing device 108 in response to the particular user query, each result being representative of one of the plurality of electronic documents. Accordingly, the user queries that have the same or similar user click data for a result such as, for example, a particular entity name can be linked to the electronic document representative of the particular entity name.

[0064] For user queries that do not include user click data, the mapping module 128 associates each of these user queries to one of the plurality of electronic documents that has the highest frequency of occurrence among the user queries with user click data. In some embodiments, the mapping module 128 associates each of the remaining user queries after the filtering to one of the plurality of electronic documents that has the highest frequency of occurrence among the user queries with user click data within the same cluster. This frequency-based mapping performed by the mapping module 128 ensures that the association of the data points without the user click data accurately reflects user click data.

[0065] According to some embodiments, the mapping module 128 can validate the mapping between the user queries and the electronic documents in the second index. In some embodiments, the mapping module 128 can validate the mapping according to the embedding vector values representative of the electronic documents and the embedding vector values representative of the user queries in the clusters. In some embodiments, the mapping module 128 can utilize a machine learning model to validate the associations between the user queries and the electronic documents. In some embodiments, the mapping module 128 can utilize a large language model to validate the associations between the user queries and the electronic documents based on the corresponding text data.

[0066] The search system 102 can include an index module 130. The index module 130 can generate a first index representative of how often specific queries are made according to a density and size of each cluster. The index module 130 can generate a second index representative of a plurality of documents based on the mapping of the user queries to the electronic documents of a third index. In this regard, the search system 102 can include a third index including a plurality of electronic documents stored in, for example, data store 104 or computer-readable memory 118. The third index can include the plurality of the electronic documents of the second index. That is, the second index includes the electronic documents from the third index (e.g., global index) that have been mapped to the user queries. In some embodiments, the second index includes the electronic documents that have been mapped to the user queries in the remaining clusters after one or more of the clusters have been filtered out.

[0067] The index module 130 generates the first index and the second index to improve the searching experience provided by the search system 102. The first index is representative of how often specific queries were made. That is, the first index is representative of a frequency that each of the plurality of documents of the second index was queried by the users according to the user queries. In some embodiments, the first index is based on the user behavior data. In some embodiments, the first index is based on the user click data. The second index can include a plurality of electronic documents from the electronic documents in the third index based on the mapping to the user queries by the mapping module 128. The first index serves as an indicator of a popularity of each of the plurality of electronic documents while the second index fills the gaps between the vocabulary used by users in the user queries and the text data representative of the electronic documents. The dual-list generation of the index module 130 enables the search system 102 to perform the search function based on user interests and relevance, which reduces a likelihood of drop-off and that the set of results provided by the search system 102 in response to the user query includes a matching result. In some embodiments, the first index can be utilized to rank the electronic documents in the second index for a search result.

[0068] The search system 102 includes a machine learning module 132. Various embodiments herein can employ artificial-intelligence, neural network models, deep learning neural network models, deep q-learning neural network models, and / or machine learning systems and techniques to facilitate training the models from scratch, training the models using audit data, training the models using reinforcement learning for continual learning, determining decisions as output predictions based on applying the input data to the models, other processes, or any combination thereof. Although the one or more embodiments are described in the present disclosure in the context of determining decisions by services at a domain in response to the context of electronic transactions, it is to be appreciated that the various embodiments can be utilized in a networked system such as, for example, system 100 for any of a plurality of purposes including, but not limited to, searching electronic documents, fulfilling transactions, authentications, content recommendations, managing threats including identifying suspicious or fraudulent transactions, learning user behavior, context-based scenarios, preferences, etc. in order to facilitate the system 100 taking automated action with high degrees of confidence for the computing devices performing transactions on the network 112. Utility-based analysis can be utilized to factor benefit of taking an action against cost of taking an incorrect action. Probabilistic or statistical-based analyses can be employed in connection with the foregoing and / or the following.

[0069] It is noted that systems and / or associated controllers, servers, or M.L. components herein such as discussed above in context of embedding module 124, cluster module 126, and mapping module 128, and the other functional modules of search system 102 in FIG. 1 can include artificial intelligence component (s) which can employ an artificial intelligence (AI) model, neural network or a neural network model, or M.L. or a M.L. model, that can learn to perform the above or below described functions (e.g., via training data and / or feedback data) . In some embodiments, the search system 102 can include a machine learning model configured to utilize natural language processing (NLP) to determine a context of a user query based on text data. In other embodiments, the search system 102 can include a machine learning model configured to utilize one or more techniques to determine a context of the user query based on text data, log data, historical data, user behavior data, user click data, image data, sequential data, other types of data, or any combination thereof. In some embodiments, the model can include, for example, a small language model, medium language model, large language model.

[0070] In some embodiments, the system 100 and / or the search system 102 can include a M.L. module including an A.I. and / or M.L. model that can be trained (e.g., via supervised and / or unsupervised techniques) to perform one or more of the above or below-described functions using training data including various context conditions that correspond to various management operations. In one example, an A.I. and / or M.L. model can further learn (e.g., via supervised and / or unsupervised techniques) to perform the above or below-described functions using training data including feedback data, where such feedback data can be collected and / or stored (e.g., in memory 118 or data store 104) by interface module 120, data module 122, or by an M.L. component of search system 102. In this example, such feedback data can include the various instructions described above / below that can be input, for instance, to a system herein, over time in response to observed / stored context-based information.

[0071] A.I.  / M.L. components herein can initiate an operation (s) associated with the one or more functional modules 120, 122, 124, 126, 128, 130 of the search system 102 based on a defined level of confidence determined using information (e.g., feedback data) . For example, based on learning to perform such functions described above using feedback data, performance information, and / or past performance information herein, an M.L. model herein can initiate an operation associated with providing one or more results such as, for example, a set of entity names, as output predictions based on the input data applied to the model including context data from the log data including, but not limited to, user data, account data, device data, location data, language preference data, application data, transaction data, historical data, inventory data, user behavior data, sequence data, other types of data at search system 102 or transaction processing system 106, or any combination thereof. In another example, based on learning to perform such functions described above using feedback data, an M.L. model can be trained from scratch, trained using reinforcement learning, trained using continual learning, or trained using data based on the activity of the services at the domain.

[0072] In an embodiment, the M.L. model can perform a utility-based analysis that factors cost of initiating the above-described operations versus benefit. In this embodiment, an artificial intelligence component can use one or more additional context conditions to determine an appropriate distance threshold or context information, or to determine an update for a tuning model.

[0073] To facilitate the above-described functions, an M.L. model herein can perform classifications, correlations, inferences, and / or expressions associated with principles of artificial intelligence. For instance, an M.L. model can employ an automatic classification system and / or an automatic classification. In one example, the M.L. model can employ a probabilistic and / or statistical-based analysis (e.g., factoring into the analysis utilities and costs) to learn and / or generate inferences. The M.L. model can employ any suitable machine-learning based techniques, statistical-based techniques and / or probabilistic-based techniques. For example, the M.L. model can employ expert systems, fuzzy logic, support vector machines (SVMs) , Hidden Markov Models (HMMs) , greedy search algorithms, rule-based systems, Bayesian models (e.g., Bayesian networks) , neural networks, other non-linear training techniques, data fusion, utility-based analytical systems, systems employing Bayesian models, and / or the like. In another example, the M.L. model can perform a set of machine-learning computations. For instance, the M.L. model can perform a set of clustering machine learning computations, a set of logistic regression machine learning computations, a set of decision tree machine learning computations, a set of random forest machine learning computations, a set of regression tree machine learning computations, a set of least square machine learning computations, a set of instance-based machine learning computations, a set of regression machine learning computations, a set of support vector regression machine learning computations, a set of k-means machine learning computations, a set of spectral clustering machine learning computations, a set of rule learning machine learning computations, a set of Bayesian machine learning computations, a set of deep Boltzmann machine computations, a set of deep belief network computations, and / or a set of different machine learning computations.

[0074] In some embodiments, the M.L. model can utilize one or more clustering techniques including, but not limited to, density-based clustering, distribution-based clustering, centroid-based clustering, hierarchical based clustering, or any combinations thereof. In addition, the one or more models can apply one or more clustering algorithms including, but not limited to, k-means clustering algorithms, density-based clustering algorithms, Gaussian mixture model algorithms, balanced iterative reducing and clustering using hierarchies (BIRCH) algorithms, propagation clustering algorithms, mean-shift clustering algorithms, order point clustering, agglomerative hierarchy clustering algorithms, other algorithms, or any combinations thereof. For example, the model can apply the one or more centroid-based clustering models to determine clusters using k-means clustering algorithms.

[0075] In addition, following each user query, the search system 102 can record the information generated from the user query as log data for subsequent analysis, thereby creating a feedback loop that continuously identifies and updates the connections between user queries and the plurality of electronic documents at regular intervals. The solution framework of the search system 102 is also capable of accommodating changes in life cycles of the external entities such as, for example, new financial institutions opening or closing. Through the user interface, the second index can also be edited by users and re-indexed within minutes. Furthermore, for cases where a new entity cannot be identified based on the log data or the indexes, the machine learning module 132 can utilize one or more A.I. or M.L. models to generate new electronic documents representative of the new entity to facilitate the linking of the user account with an external account in the external computing network of the external entity.

[0076] It is to be appreciated that the system 100 and / or the search system 102 is not intended to be limited to include functional modules 120, 122, 124, 126, 128, 130, 132. That is, the search system 102 can include any one of these functional modules or other functional modules so as to able to perform the operations in accordance with the embodiments of the present disclosure.

[0077] The present disclosure describes the search system 102 in the context of electronic transactions performed utilizing the transaction processing system 106 in system 100. It is to be appreciated that the search system 102 can be utilized by one or more systems associated with an entity. In some embodiments, the system 100 can include one or more systems including the transaction processing system 106, and the one or more systems can utilize the search system 102 for their respective process flows. In some embodiments, the system 100 associated with an entity can include the transaction processing system 106, and one or more other systems associated with the entity can be in electronic communication with the search system 102 in system 100 using the network 112 or some other network. In some embodiments, the system 100 can include the one or more other systems. In some embodiments, the one or more other systems can be external to system 100. According to some embodiments, the search system 102 can be utilized in the context of a digital wallet 140 (FIG. 11) . Each digital wallet 140 can be associated with a respective user. The digital wallet 140 can be utilized by the user to perform one or more process flows associated with goods or services offered by the entity in, for example, the system 100. For example, the search system 102 can be implemented by a process flow via the digital wallet 140 for a user associated with the user computing device 108 can include an application for implementing a process flow to link their user account with an external account of an external system to perform electronic transactions with other users using the transaction processing system 106 of system 100. For example, the search system 102 can be implemented by a process flow via the digital wallet 140 for a user associated with user computing device 108 to link their user account with an external account of an external system to add electronic funds from the externa account to the user account in the system 100.

[0078] FIG. 2 is a flow diagram of an example method 200 for providing a set of results in response to a user query, according to some embodiments. The method 200, or one or more portions of the method 200, can be performed by the search system 102 in conjunction with the transaction processing system 106, and thus can be computer-implemented. FIG. 3 is a block diagram of an example system 300 for performing the method 200 of FIG. 2, according to some embodiments. The method 200 will be described in conjunction with the system 300. According to some embodiments, the method 200 can be performed in on offline mode.

[0079] At 202, the method 200 can include obtaining a text input for user queries from a plurality of users. In some embodiments, each of the user queries corresponds to a request for one of a plurality of electronic documents. In some embodiments, the user queries can be historical user queries obtained from a data store such as, for example, data store 104 in FIG. 1. In some embodiments, the user query can be obtained as further described with regards to interface module 120 in FIG. 1. In FIG. 3, the user queries is shown as query 302. For example, a user can initiate a process flow for adding electronic funds to the user’s account in a networked computing system associated with an entity from an external account, and the process flow can implement the search system to receive the user query through an electronic user interface.

[0080] At 204, the method 200 can include retrieving a first log data based on the user queries. In some embodiments, the first log data includes user behavior data for the plurality of users associated with the user queries. In some embodiments, the user behavior data can be historical user behavioral data. In some embodiments, the user behavior data can include user click data corresponding to a user selection at a user interface of a result provided to the user interface in response to the user query. In some embodiments, the log data can be retrieved as further described with regards to data module 122 in FIG. 1. In FIG. 3, the log data is shown as log data 304. For example, the process flow for adding electronic funds to the user’s account from the external account can implement the search system to retrieve log data for users performing similar operations using the networked computing system associated with the entity.

[0081] According to some embodiments, the method 200 can further include retrieving a plurality of electronic documents. In some embodiments, the method can further include retrieving the plurality of electronic documents corresponding to the second index from a third index. In some embodiments, each of the plurality of electronics documents includes text data representative of at least one of an entity name, a location, a unique identifier, or one or more aliases. For example, in some embodiments, an electronic document can include one or more sets of text string data representative of one or more aliases of an entity name based on historical user queries by a plurality of users. In some embodiments, the electronic documents can be retrieved as further described with regards to data module 122 in FIG. 1. In FIG. 3, the electronic documents is shown as documents 306.

[0082] At 206, the method 200 can include using a machine learning model and determining embedding vectors representative of the user queries based on the text input and the first log data. In some embodiments, the embedding vectors are representative of semantic characteristics according to the text input and representative of contextual relationships according to the first log data. In some embodiments, the first log data can include data associated with the plurality of users including at least one of account data, application data, location data, or language preference data.

[0083] The embedding vectors can be transformed from the text input data using the machine learning model. The embedding vectors can also be transformed from the log data using the machine learning models and associated with the corresponding user queries. In some embodiments, the machine learning models can include an embedding model as further described with regards to embedding module 124 in FIG. 1. In FIG. 3, the embedding vectors are shown as embeddings 308.

[0084] At 208, the method 200 can include partitioning the embedding vectors into a respective one of a plurality of clusters according to characteristics of each of the embedding vectors. In some embodiments, each cluster can include one or more of the user queries received as input. In some embodiments, the user queries can be partitioned in the embedding space according to the embedding vector values associated therewith, with user queries having similar embedding vector values being located near each other relative user queries having different embedding vector values. The partitioning of the embedding vectors can be performed using the machine learning model. In some embodiments, the machine learning model can include a clustering model. In some embodiments, the machine learning model can include a clustering model as further described with regards to cluster module 126 in FIG. 1. In FIG. 3, the clustering of the embedding vectors is shown at cluster 310.

[0085] At 210, the method 200 can include mapping each of the plurality of clusters to a respective one of the plurality of electronic documents according to the partitioned embedding vectors associated with each cluster. In some embodiments, the mapping can include associating, for each user query in the cluster, the user query to one of the plurality of electronic documents of the second index according to the embedding vectors values. In some embodiments, the mapping can be based on the user behavior data. In some embodiments, the mapping can be based on the user behavior data associated with other user queries nearest the user query. The user behavior data can include user click data representative of a corresponding user’s selection of a result at the user interface in response to the user’s query. In some embodiments, the user queries can be mapped to the electronic documents as further described with regards to mapping module 128 in FIG. 1. In FIG. 3, the mapping is shown at mapping 312.

[0086] At 212, the method 200 can include generating a first index representative of a frequency each of the plurality of electronic documents is queried by the plurality of users according to the partitioned embedding vectors and generating a second index representative of the plurality of electronic documents according to the mapping. The second index can include a plurality of the electronic documents from the third index according to the mapping. That is, the electronic documents in the second index can be from the third index that have been mapped to the user queries. In this regard, in some embodiments, the second index can include at least a subset of the electronic documents of the third index. In some embodiments, the second index includes the electronic documents that have been mapped to the remaining user queries in the clusters.

[0087] The first index and the second index improves the searching experience. The second index can include a plurality of the electronic documents retrieved based on the mapping from all of the electronic documents in the third index. The first index serves as an indicator of a popularity of each of the plurality of electronic documents while the second index fills the gaps between the vocabulary used by users and the electronic documents. The dual-list generation enables performing the search function of method 200 based on user interests and relevance, which reduces a likelihood of drop-off and that the provided set of results in response to a subsequent user query includes a relevant, matching result. In some embodiments, the first index can be utilized to rank the electronic documents in the second index for a search result. In some embodiments, the generating of the first index and the second index can be performed as further described with regards to index module 130 in FIG. 1. In FIG. 3, the first index is shown at query index 314 and the second index is shown at document index 316.

[0088] At 214, the method 200 can include storing the first index and the second index for fulfilling subsequent user queries corresponding to requests for return of at least one of the plurality of electronic documents from users. The first index and the second index can be stored in a data store. In some embodiments, the first index and the second index can be stored in a same data store as the log data. In some embodiments, the first index and the second index can be stored in a same data store as the electronic documents. In some embodiments, the storing of the first index and the second index can be performed as further described with regards to data module 122 in FIG. 1.

[0089] FIG. 4 is a flow diagram of an example method 400 for providing a set of results in response to a user query, according to some embodiments. The method 400, or one or more portions of the method 400, can be performed by the search system 102 in conjunction with the transaction processing system 106, and thus can be computer-implemented. The method 400 can be an embodiment of operations 204, 206, 208 of method 200, according to some embodiments. The method 400 will be described in conjunction with the system 300 of FIG. 3. According to some embodiments, the method 400 can be performed in on offline mode.

[0090] At 402, the method 400 can include using the machine learning model and determining embedding vectors representative of the plurality of electronic documents. The text data of the plurality of electronic documents can be transformed using the machine learning model into embedding vector values. In some embodiments, the embedding vector values of the electronic documents can be provided as further described with regards to embedding module 124 in FIG. 1. In FIG. 3, the embeddings of the electronic documents is shown at embeddings 308. For example, the process flow for adding electronic funds to the user’s account from the external account can implement the search system to determine the relevant embeddings from the plurality of electronic documents in the context of adding electronic funds to user accounts by the users.

[0091] At 404, the method 400 can include using the machine learning model and validating the mapping between the cluster and the plurality of electronic documents. In some embodiments, the validation is based on the embedding vectors values representative of the user queries in the cluster and the embedding vectors values representative of the plurality of electronic documents. In some embodiments, the validating of the mapping can be performed as further described with regards to mapping module 128 in FIG. 1. In FIG. 3, the validating can be performed at mapping 312.

[0092] According to some embodiments, in method 400, the validating of the mapping between the user queries and the electronic documents in the second index can be according to the embedding vector values representative of the electronic documents and the embedding vector values representative of the user queries in the clusters. In some embodiments, the validating of the mapping between the user queries and the electronic documents in the second index can include utilizing a machine learning model to validate the associations between the user queries and the electronic documents. In some embodiments, the machine learning model includes a language model. The machine learning model can utilize the language model to validate the mapping between the text input data of the user queries and the text data of the electronic documents by determining a semantic relationship therebetween using the embedding vectors representative of the user queries and the embeddings vectors representative of the electronic documents. In some embodiments, the language model can be a large language model. In some embodiments, the language model can be a medium language model. In some embodiments, the language model can be a small language model.

[0093] FIG. 5 is a flow diagram of an example method 500 for refining clusters representative of the user queries, according to some embodiments. The method 500, or one or more portions of the method 500, can be performed by the search system 102 in conjunction with the transaction processing system 106, and thus may be computer-implemented. The method 500 can be an embodiment of operations 208, 210, 212 of the method 200 of FIG. 2, according to some embodiments.

[0094] At 502, the method 500 can include identifying one or more clusters of the plurality of clusters based on the embedding vectors being less than a predefined threshold. In some embodiments, the identified clusters can be representative of low-intensity clusters. In some embodiments, the low-intensity clusters can include outlier user queries. In some embodiments, the low-intensity clusters can include random click data for a set of results provided to the corresponding user in response to the user query. In some embodiments, the identifying of the one or more clusters of the plurality of clusters can be performed as further described with regards to mapping module 128 of FIG. 1. For example, a process flow for transferring electronic funds from one user to another user of the networked computing system associated with the entity can identify one or more clusters of the plurality of clusters based on user click data associated with users performing similar transactions in the networked computing system. For example, a particular process flow can utilize the search system to identify the one or more clusters of the plurality of clusters having a low-intensity of vectors representative of user click data associated with the particular process flow that is below the predefined threshold.

[0095] At 504, the method 500 can include refining a search space by filtering the identified one or more clusters from the plurality of clusters. In some embodiments, the mapping of the plurality of clusters to the electronic documents of the second index can include mapping each of the remaining clusters of the plurality of clusters that were not filtered out from the embeddings space to a respective one of the plurality of electronic documents. In this regard, the user queries in these low-intensity clusters can be filtered out from the embeddings space by the machine learning model to improve the quality of the data and so that the remaining user queries that remain in the embeddings space are representative of relevant and significant user interactions that can be used for further analysis. For example, to rank the electronic documents of the second index according to the personal data of a given user and to provide at least a subset of the electronic documents of the second index in response to a user query from the given user. In some embodiments, the filtering of the identified one or more clusters can be performed as further described with regards to mapping module 128 of FIG. 1.

[0096] FIG. 6 is a flow diagram of an example method 600 for mapping the clusters of user queries to an index, according to some embodiments. The method 600, or one or more portions of the method 600, can be performed by the search system 102 in conjunction with the transaction processing system 106, and thus may be computer-implemented. The method 600 can be an embodiment of operations 208, 210 212 of the method 200 of FIG. 2, according to some embodiments. FIG. 7 is a block diagram of an example system 700 for mapping the clusters of user queries to an index, according to some embodiments. The method 600 will be described in conjunction with the system 700 of FIG. 7.

[0097] At 602, the method 600 can include associating, for each user query in the cluster including the user behavior data, the user query to one of a plurality of electronic documents of a third index according to the embedding vectors values. In some embodiments, the third index can correspond to a global index including a plurality of electronic documents. The electronic documents in the third index can be representative of, for example and without limitation, entities, external entities, policies, documents, user accounts, products including goods or services, etc. In some embodiments, the third index can include the plurality of electronic documents of the second index. In some embodiments, based on the association (i.e., linking) of the user query to one of the plurality of electronic documents in the third index, the electronic document can be provided in the second index.

[0098] In a cluster, each user query can include embedding vector values representative of the user behavior data. In some embodiments, the user behavior data can correspond to user click data (e.g., historical user click data) representative of the user’s interaction with one of a set of results representative of a set of electronic documents in the second index that can be provided to the user via the user interface in response to the user query. For example, the set of results can be a set of text string data (e.g., entity names) representative of a ranked subset of electronic documents from the second index in response to a user query. In some embodiments, the associating of the user query that includes the user behavior data to one of the electronic documents of the third index can be performed as further described with regards to mapping module 128 in FIG. 1. In FIG. 7, the user queries including the user behavior data is shown as queries 706, the documents with user interactions is shown as documents 710, the user queries mapped to electronic documents is shown as mapped queries 714.

[0099] At 604, the method 600 can include associating, for each user query in the cluster not including the user behavior data, the user query to one of the plurality of electronic documents of the third index based on a frequency of the association of the user queries including the user behavior data. A user query may not include user behavior data for any of a plurality of reasons. For example, in some embodiments, a user query in a cluster may not include the corresponding user behavior data due to the set of results provided to the user interface not including a correct result. For example, in some embodiments, a user query in a cluster may not include the corresponding user behavior data because the user did not select one of the set of results at the user interface or abandoned the query without providing the user interaction. For example, the process flow implementing the search system to provide the set of results in response to a user search query may be a new application provided by the entity that does not include previous user behavior data.

[0100] For a user query that does not include the relevant user behavior data, the machine learning model may associate the user query to one of the plurality of electronic documents having the highest frequency of interactions of the other nearest user queries. In some embodiments, for each user query in the cluster not including the user behavior data, the user query can be associated with one of the plurality of electronic documents of the third index based on a frequency of the association of the other user queries including the user behavior data in the same cluster. In some embodiments, for each user query in the cluster not including the user behavior data, the user query can be associated with one of the plurality of electronic documents of the third index based on a frequency of the association of the other user queries including the user behavior data in one or more clusters nearest the user query. In some embodiments, the associating of the user query not including the user behavior data to the electronic documents of the third index can be performed as further described with regards to mapping module 128 of FIG. 1. In FIG. 7, the queries not including the user behavior data is shown as queries 708, the documents with the highest frequency of associations with the user queries including the user behavior data is shown as documents 712. In FIG. 7, the mapped user queries is shown as mapped queries 714. In some embodiments, the mapped queries 714 can include queries 706 and queries 708. In some embodiments, the mapped queries 714 can include the electronic documents that are associated with the queries 706 and queries 708 from the third index. In FIG. 7, the validating of the mapped queries 714 is shown at validate 716, which can be performed as described with regards to mapping module 128 in FIG. 1 and with regards to method 400 of FIG. 4.

[0101] FIG. 8 is a flow diagram of an example method 800 for providing a set of results in response to a user query, according to some embodiments. The method 800, or one or more portions of the method 800, can be performed by the search system 102 in conjunction with the transaction processing system 106, and thus may be computer-implemented. The method 800 can be an embodiment of method 200 of FIG. 2, according to some embodiments. FIG. 9 is a block diagram of an example system 900 for providing the set of results in response to the user query, according to some embodiments. The method 800 will be described in conjunction with the system 900 of FIG. 9. According to some embodiments, the method 800 can be performed in on online mode.

[0102] At 802, the method 800 can include receiving a text input of a given user query from a user computing device of a given user. In some embodiments, the user can provide the user query as text input data at a user interface at the user computing device. The user query can be received in an online mode. That is, in some embodiments, the user query can be received at a time period after the first index and the second index are generated and stored, as described with regards to method 200 in FIG. 2. In some embodiments, the receiving of the text input can be performed as described with regards to interface module 120 in FIG. 1. In FIG. 9, the user query is shown as user query 902.

[0103] At 804, the method 800 can include retrieving, in response to the given user query, a second log data comprising user behavior data associated with the given user. In some embodiments, the second log data includes data associated with the given user. The second log data can include at least one of account data, application data, location data, or language preference data. The account data can be associated with the user providing the user query. The application data can be the applications stored in the user computing device. For example, the applications stored in the user computing device can be associated with external computing networks of external entities. The location data can be a location of the user. For example, the location data can be provided by the user at the user interface as input. For example, the account data can include the location data defined by the user when setting up the user’s account. In some embodiments, the location data can be a location of the user computing device. In some embodiments, the location data can be representative of a state. In some embodiments, the location data can be representative of a city. In some embodiments, the location data can be representative of a country. The language preference data can be defined by the user. The language preference data can include one or more languages. For example, the language preference data can be input by the user at the user interface. For example, the account data can include the language preference data defined by the user when setting up the user’s account. In FIG. 9, the log data is shown as log data 904. In some embodiments, the retrieving of the second log data can be performed as described with regards to data module 122 in FIG. 1. In FIG. 9, the log data 904 can include behavior data 906, account data 908, location data 910, languages data 912, or any combinations thereof, according to some embodiments.

[0104] According to some embodiments, the method 800 includes configuring a processing pipeline to fulfill the user query. The processing pipeline can be an elastic search pipeline configured to receive the user query and retrieve the log data. The processing pipeline can also be configured to retrieve the first index and the second index, and the processing pipeline can be configured to utilize the first index and the second index to provide the set of results to the user in response to the user query, as described with regards to interface module 120, data module 122, embedding module 124, cluster module 126, mapping module 128, index module 130, and machine learning module 132 in FIG. 1, and as will be further described herein. In some embodiments, the processing pipeline can perform operations as described with regards to embedding module 124, cluster module 126, mapping module 128, index module 130, or any combination thereof, of FIG. 1. In FIG. 9, the processing pipeline is shown as pipeline 914, the first index is shown as first index 916, the second index is shown as second index 918. As used herein, the term “elastic search pipeline” refers to a series of processing tasks to transform data before being utilized in an index or to generate an index. For example, the process flow can implement the pipeline 914, and the pipeline 914 can include indexing and searching pipelines including an online search phase where the user search query is combined with personal data of the user including, but not limited to, a list of external entity applications installed in the user’s computing device, systems associated with external entities that have already been linked to other users in the system associated with the entity, and other user data including, but not limited to, the user’s country, state, and language preference data, such that the pipeline 914 can retrieve the appropriate index and display the ranked set of results at the electronic user interface for selection by the user.

[0105] At 806, the method 800 can include using the machine learning model and ranking the plurality of electronic documents of the second index based on the first index and the second log data. In some embodiments, the second index includes the plurality of electronic documents mapped to the user queries (i.e., historical user queries) as described with regards to index module 130 in FIG. 1 and / or method 200 in FIG. 2.

[0106] According to some embodiments, the plurality of electronic documents in the second index can be ranked according to the first index, the second index, or both. In some embodiments, the plurality of electronic documents in the second index can be ranked according to the frequency of specific queries from the first index. For example, in some embodiments, the electronic document having the highest frequency of user queries as described with regards to method 200 can have the highest ranking in the second index. For example, in some embodiments, each electronic document in the second index can include a ranking score associated therewith according to a frequency each electronic document was queried by the user queries in the offline mode as described with regards to method 200. In some embodiments, the plurality of electronic documents in the second index can also be ranked according to second log data of the user. In some embodiments, the plurality of electronic documents in the second index can initially be ranked according to first index, and the second log data can then be utilized to re-rank the plurality of electronic documents in the second index. For example, in some embodiments, the electronic document associated with an entity can be given a higher ranking in the second index based on the user computing device having an application associated with the entity stored in the user computing device. For example, in some embodiments, an electronic document can be given a higher ranking in the second index than another electronic document based on the entity associated with the electronic document having location data representative of a brick-and-mortar location near a location of the user and the other electronic document not having location data representative of a brick-and-mortar location near the location of the user. For example, in some embodiments, each electronic document in the second index can include a ranking score associated therewith according to a frequency each electronic document was queried by the user queries in the offline mode according to the first index and according to the second log data. In FIG. 9, the ranking is shown as document ranking 920.

[0107] At 808, the method 800 can include using the machine learning model and providing a subset of the ranked plurality of electronic documents from the second index to the given user. In some embodiments, the subset of electronic documents can be provided to the user as a set of text string data, each text string representative of one of the subset of electronic documents. In some embodiments, the subset of the plurality of electronic documents of the second index are provided to the given user based on the machine learning model using at least one of: a fuzzy search algorithm, a prefix search algorithm, or a precise search algorithm. The subset of the ranked plurality of electronic documents from the second index can be provided to the user as described with regards to index module 130 and interface module 120 in FIG. 1.

[0108] According to some embodiments, the method 800 can include receiving user behavior data in response to the subset of the ranked plurality of electronic documents provided to the user computing device as a set of results. The user behavior data can be feedback data indicative of the user’s selection of one of the set of results provided to the user. For example, the set of results can be the set of text strings representative of a set of entities provided to the user computing device via the user interface, and the user behavior data can be user click data from the user’s interaction with one of the set of text strings at the user interface or from the user’s interaction with a field representative of one of the set of text strings. In FIG. 9, the user behavior data is shown as user behavior 922.

[0109] At 810, the method 800 can include storing a user interaction with the subset of the ranked plurality of electronic documents as feedback data for fulfilling the subsequent user queries. In some embodiments, the method 800 can include storing the user interaction as search log data for fulfilling the subsequent user queries. In FIG. 9, the search log data is shown as search logs 924. The data store that stores the search log data is shown as data store 926.

[0110] According to some embodiments, the method 800 can further include operations similar to method 200 in FIG. 2 including, but not limited to, transforming the user query and the user interaction data into embedding vectors, allocating the embeddings vectors representative of the user query and the user interaction data into the partitioned embeddings space along with the other user queries (i.e., historical user queries) , mapping the user query to the electronic documents according to the user behavior data or according to the user behavior data of the other nearest user queries, and generating a first index representative of a frequency each of the plurality of electronic documents is queried by the users including the given user according to the partitioned embedding vectors and a second index representative of the plurality of electronic documents according to the mapping, and storing the updated first index and the updated second index in a data store for fulfilling subsequent user queries. For example, the updated first index and the second index can be generated using the given user query during a next scheduled offline mode.

[0111] According to some embodiments, the methods 200, 400, 500, 600, 800 can further include generating a third index. In some embodiments, generating the third index can include updating the electronic documents included in the third index with a new document. In some embodiments, generating the third index can include updating the electronic documents included in the third index by removing an electronic document. For example, in some embodiments, the third index that is stored in a data store of a system associated with an entity such as, for example, system 100 of FIG. 1, can include electronic documents representative of external computing networks associated with external entities, and users performing electronic transactions on the system can also have accounts with these external computing networks that they want to link to their user account in the system of the entity. In this regard, as new external computing networks are created or become unavailable, the third index may need to be updated to ensure accuracy and so that the results provided to users in response to their user queries includes any new results that the user may be searching for. For example, the third index may be representative of external computing networks of different financial institutions and may need to be updated whenever a new financial institution is created so the users of the system 100 of FIG. 1 can search and add an account with the external computing network of the new financial institution. In FIG. 9, the third index is shown as 9320, the generating of the documents is shown as generate 932, the updating of the documents is shown as edit 934, the new document is shown as new document 936, and the document being removed is shown as close document 938.

[0112] According to some embodiments, generating the third index can include editing an electronic document stored in the third index. In some embodiments, editing an electronic document stored in the third index can include updating the data of the electronic document. In this regard, each electronic document can include a plurality of data including text data. The electronic document can include, for example and without limitation, a name data or label data, location data, unique identifier data, alias data, or any combination thereof. For example, in some embodiments, the name data can be representative of an entity name. For example, in some embodiments, the name data can be representative of a product name. For example, in some embodiments, the location data can be representative of a location of the entity. For example, in some embodiments, the location data can be representative of one or more locations of the entity. For example, in some embodiments, the unique identifier data can be a routing number. For example, in some embodiments, the unique identifier data can be a SWIFT code. For example, in some embodiments, the one or more aliases from the user queries can be associated with the electronic document.

[0113] In some embodiments, updating the electronic document can include using the machine learning model to update an alias that is associated with the electronic document. In some embodiments, one or more user queries that is associated with an electronic document (e.g., based on user behavior data) can include a common text input. Although this common text input may not be an exact match to the name data or to a portion of the name data representative of the electronic document such as, for example, due to a vocabulary mismatch, the electronic document data can be updated to include this common text input as an alias using the machine learning model based on the context. In some embodiments, updating the electronic document can include using the machine learning model to update a location data that is associated with the electronic document. In some embodiments, the location data can be retrieved from, for example, the external computing network of the external entity using the machine learning model. In this regard, the electronic document data can be updated using the machine learning model to improve the generating of the first index and the second index in the offline mode and to improve the results returned to a user in response to a user query in the online mode.

[0114] According to some embodiments, one or more of the methods 200, 400, 500, 600, 800 may be embodiments for providing a set of results utilizing a search system such as, for example, the search system 102 in system 100, by one or more process flows. For example, a system associated with an entity can provide a digital wallet including one or more applications representative of various goods and services associated with the entity to each user through an electronic user interface, and the user can interact with an application at the digital wallet through the electronic user interface to implement a process flow for a corresponding goods and services. In addition, through the electronic user interface, when the user implements the process flow by interacting with the application in the digital wallet, the process flow can include performing the method 200 to provide the set of results to the user through the electronic user interface in response to the user query. For example, the electronic user interface can include a search portion configured to receive the user query from the user as a text input, and the set of results can be displayed to the user at the search portion of the electronic user interface in response to the user query. For example, the set of results can be generated as one or more selectable fields adjacent (e.g., below) a search field for receiving the user query input. In some embodiments, the system 100 can include one or more systems that can implement the one or more process flows that utilizes the search system 102. For example, the transaction processing system 106 can be utilized to implement a process flow using the search system 102. In some embodiments, the system 100 can include the transaction processing system 106, and one or more other systems external to system 100 can be in electronic communication with the system 100 and search system 102 through network 112 or some other network such that the system 100 and the one or more other systems can implement the one or more process flows that utilizes the search system 102.

[0115] FIG. 10 is a block diagram of an example computing network device 1000, according to some embodiments.

[0116] The computing system 1000 can be, for example, a desktop computer, laptop, smartphone, tablet, or any other such device having the ability to execute instructions, such as those stored within a non-transient, computer-readable medium. Furthermore, while described and illustrated in the context of a single computing system 1000, those skilled in the art will also appreciate that the various tasks described hereinafter can be practiced in a distributed environment having multiple computing systems 1000 linked via a local or wide-area network in which the executable instructions can be associated with and / or executed by one or more of multiple computing systems 1000.

[0117] In its most basic configuration, computing system environment 1000 typically includes at least one processing unit 1002 and at least one memory 1004, which can be linked via a bus 1006. Depending on the exact configuration and type of computing system environment, memory 1004 can be volatile (such as RAM 1010) , non-volatile (such as ROM 1008, flash memory, etc. ) or some combination of the two. Computing system environment 1000 can have additional features and / or functionality. For example, computing system environment 1000 can also include additional storage (removable and / or non-removable) including, but not limited to, magnetic or optical disks, tape drives and / or flash drives. Such additional memory devices can be made accessible to the computing system environment 1000 by means of, for example, a hard disk drive interface 1012, a magnetic disk drive interface 1014, and / or an optical disk drive interface 1016. As will be understood, these devices, which would be linked to the system bus 1006, respectively, allow for reading from and writing to a hard disk 1018, reading from or writing to a removable magnetic disk 1020, and / or for reading from or writing to a removable optical disk 1022, such as a CD / DVD ROM or other optical media. The drive interfaces and their associated computer-readable media allow for the nonvolatile storage of computer readable instructions, data structures, program modules and other data for the computing system environment 1000. Those skilled in the art will further appreciate that other types of computer readable media that can store data can be used for this same purpose. Examples of such media devices include, but are not limited to, magnetic cassettes, flash memory cards, digital videodisks, Bernoulli cartridges, random access memories, nano-drives, memory sticks, other read / write and / or read-only memories and / or any other method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Any such computer storage media can be part of computing system environment 1000.

[0118] A number of program modules can be stored in one or more of the memory / media devices. For example, a basic input / output system (BIOS) 1024, containing the basic routines that help to transfer information between elements within the computing system environment 1000, such as during start-up, can be stored in ROM 1008. Similarly, RAM 1010, hard drive 1018, and / or peripheral memory devices can be used to store computer executable instructions comprising an operating system 1026, one or more applications programs 1028, other program modules 1030, and / or program data 1032. Still further, computer-executable instructions can be downloaded to the computing environment 1000 as needed, for example, via a network connection. The applications programs 1028 can include, for example, a browser, including a particular browser application and version, which browser application and version can be relevant to determinations of correspondence between communications and user URL requests, as described herein. Similarly, the operating system 1026 and its version can be relevant to determinations of correspondence between communications and user URL requests, as described herein.

[0119] An end-user can enter commands and information into the computing system environment 1000 through input devices such as a keyboard 1034 and / or a pointing device 1036. While not illustrated, other input devices can include a display, microphone, a joystick, a game pad, a scanner, etc. These and other input devices would typically be connected to the processing unit 1002 by means of a peripheral interface 1038 which, in turn, would be coupled to bus 1006. Input devices can be directly or indirectly connected to processor 1002 via interfaces such as, for example, a parallel port, game port, firewire, or a universal serial bus (USB) . To view information from the computing system environment 1000, a monitor 1040 or other type of display device can also be connected to bus 1006 via an interface, such as via video adapter 1033. In addition to the monitor 1040, the computing system environment 1000 can also include other peripheral output devices, not shown, such as speakers and printers.

[0120] The computing system environment 1000 can also utilize logical connections to one or more computing system environments. Communications between the computing system environment 1000 and the remote computing system environment can be exchanged via a further processing device, such a network router 1048, that is responsible for network routing. Communications with the network router 1048 can be performed via a network interface component 1044. Thus, within such a networked environment, e.g., the Internet, World Wide Web, LAN, or other like type of wired or wireless network, it will be appreciated that program modules depicted relative to the computing system environment 1000, or portions thereof, can be stored in the memory storage device (s) of the computing system environment 1000.

[0121] The computing system environment 1000 can also include localization hardware 1046 for determining a location of the computing system environment 1000. In embodiments, the localization hardware 1046 can include, for example only, a GPS antenna, an RFID chip or reader, a WiFi antenna, or other computing hardware that can be used to capture or transmit signals that can be used to determine the location of the computing system environment 1000. Data from the localization hardware 1046 can be included in a callback request or other user computing device metadata in the methods of this disclosure.

[0122] The computing system, or one or more portions thereof, can embody a user computing device 108, in some embodiments. Additionally, or alternatively, some components of the computing system 1000 can embody the stand-in system 102 and / or transaction processing system 106. For example, one or more of the functional modules 120, 122, 124, 126, 128, 130, 132 can be embodied as program modules 1030. For example, the interface module 120 can be embodied as program modules 1030. In another example, the data module 122 can be embodied as program modules 1030. Some components of the computing system 1000 can embody system 100, 700, 900. For example, the pipeline 914 can be embodied as program modules 1030.

[0123] FIG. 11 is a block diagram of the example system 100, according to some embodiments.

[0124] The system 100 can provide a digital wallet 140 to each user associated with the user computing devices 108, the user computing devices 110, or both (shown as user device 108) . The digital wallet 140 can include one or more applications 142 (shown as applications 142a, 142b, 142c, through 142n) for implementing process flows 150 (shown as process flows 150a, 150b, 150c, through 150n) , which are associated with goods or services offered by the entity. The one or more applications 142 can include, but is not limited to, onboarding, peer-to-peer transactions, deposits or add funds, withdrawals, cryptocurrency-based transactions, savings plans, installment plans, other types of electronic transactions, or any combination thereof. In some embodiments, the process flow 150 associated with one or more of the applications 142 can be performed by transaction processing system 106. In some embodiments, one of the applications such as, for example, application 142a can be performed by the transaction processing system 106, and the other applications (e.g., 142b, 142c, through 142n) can be performed by one or more other systems other than the transaction processing system 106. For example, each application can be performed by a respective system of one or more systems including the transaction processing system 106.

[0125] The digital wallet and the applications 142 can be displayed, for example, to the user computing device 108 through an electronic user interface (not shown) . The electronic user interface can include one or more selectable elements, each element corresponding to one of the applications. To implement a particular process flow 150, the user can interact with the selectable element representative of the corresponding application (e.g., 142a, 142b, 142c, through 142n) . In some embodiments, the electronic user interface can be utilized to provide the user with one or more input elements generated based on the relevant portion of the process flow. For example, the process flow for an onboarding application can provide the user associated with the user computing device 108 with a popup element including a search input field provided through the user interface to receive the user query input from the user. For example, the process flow for a cryptocurrency-based transaction application can direct the user associated with the user computing device 108 to a search page including a search input field through the user interface to receive the user query input from the user query.

[0126] In some embodiments, a computer-implemented method includes: obtaining a text input for user queries from a plurality of users, each of the user queries corresponding to a request for one of a plurality of electronic documents; retrieving a first log data based on the user queries, the first log data including user behavior data associated with the plurality of users; using a machine learning model and determining embedding vectors representative of the user queries based on the text input and the first log data; partitioning the embedding vectors into a respective one of a plurality of clusters according to characteristics of each of the embedding vectors, each cluster including one or more of the user queries; mapping each of the plurality of clusters to a respective one of the plurality of electronic documents according to the partitioned embedding vectors associated with each cluster; generating a first index representative of a frequency each of the plurality of electronic documents is queried by the plurality of users according to the partitioned embedding vectors and a second index representative of the plurality of electronic documents according to the mapping; and storing the first index and the second index for fulfilling subsequent user queries corresponding to requests for return of at least one of the plurality of electronic documents from users.

[0127] In some embodiments, the computer-implemented method further includes: using the machine learning model and determining embedding vectors representative of the plurality of electronic documents; and using the machine learning model and validating the mapping between the cluster and the plurality of electronic documents. In some embodiments, the validation is based on the embedding vectors representative of the user queries in the cluster and the embedding vectors representative of the plurality of electronic documents.

[0128] In some embodiments, the computer-implemented method further includes: identifying one or more clusters of the plurality of clusters based on the embedding vectors being less than a predefined threshold; and refining a search space by filtering the identified one or more clusters from the plurality of clusters. In some embodiments, the mapping includes mapping each of the remaining clusters of the plurality of clusters to a respective one of the plurality of electronic documents.

[0129] In some embodiments, mapping each of the plurality of clusters to the respective one of the plurality of electronic documents according to the partitioned embedding vectors associated with each cluster further includes: associating, for each user query in the cluster including the user behavior data, the user query to one of a plurality of electronic documents of a third index according to the embedding vectors; and associating, for each user query in the cluster not including the user behavior data, the user query to one of the plurality of electronic documents of the third index based on a frequency of the associations of the user queries including the user behavior data. In some embodiments, the third index includes the plurality of electronic documents of the second index. In some embodiments, the second index and the third index are generated based on the associations.

[0130] In some embodiments, the computer-implemented method further includes: receiving a text input of a given user query from a user computing device of a given user; retrieving, in response to the given user query, a second log data including user behavior data associated with the given user; using the machine learning model and ranking the plurality of electronic documents of the second index based on the first index and the second log data; using the machine learning model and providing a subset of the ranked plurality of electronic documents from the second index to the given user; and storing the subset of the ranked plurality of electronic documents for fulfilling the subsequent user queries.

[0131] In some embodiments, the second log data includes data associated with the given user including at least one of: account data, application data, location data, or language preference data.

[0132] In some embodiments, the subset of the plurality of electronic documents of the second index are provided to the given user based on the machine learning model using at least one of: a fuzzy search algorithm, a prefix search algorithm, or a precise search algorithm.

[0133] In some embodiments, the embedding vectors are representative of semantic characteristics of the text input and representative of contextual relationships of the first log data.

[0134] In some embodiments, the first log data includes data associated with the plurality of users including at least one of: account data, application data, location data, or language preference data.

[0135] In some embodiments, each of the plurality of electronics documents includes text data including at least one of: an entity name, a location, a unique identifier, or one or more aliases. In some embodiments, the one or more aliases are representative of the entity name based on the user queries by the plurality of users.

[0136] In some embodiments, a system includes: a processor, and a non-transitory computer readable medium having stored thereon instructions that are executable by the processor to perform operations including: obtain a text input for user queries from a plurality of users, each of the user queries corresponding to a request for one of a plurality of entities; retrieve a first log data based on the user queries, the first log data including user behavior data associated with the plurality of users; use a machine learning model and generate a first index representative of a frequency each of the plurality of entities is queried by the plurality of users according to a clustering of the user queries and a second index representative of a mapping of the clusters of the user queries to the plurality of entities; and store the first index and the second index for fulfilling subsequent user queries corresponding to requests for return of at least one of the plurality of entities from users. In some embodiments, fulfilling the subsequent user queries includes: receive a text input of a given user query from a user computing device of a given user; retrieve, in response to the given user query, a second log data including user behavior data associated with the given user; using the machine learning model and rank the plurality of entities of the second index based on the first index and the second log data; and using the machine learning model and provide a subset of the ranked plurality of entities from the second index to the given user.

[0137] In some embodiments, the operations further including: using a machine learning model and determine embedding vectors representative of the user queries based on the text input and the first log data; partition the embedding vectors into a respective one of a plurality of clusters according to characteristics of each of the embedding vectors, each cluster including one or more of the user queries; identify one or more clusters of the plurality of clusters based on the embedding vectors representative of the user behavior data in the one or more clusters being less than a predefined threshold; refine a search space by filtering the identified one or more clusters from the plurality of clusters; and mapping each of the remaining plurality of clusters to a respective one of the plurality of entities according to the partitioned embedding vectors associated with each cluster.

[0138] In some embodiments, mapping each of the remaining plurality of clusters to the respective one of the plurality of entities according to the partitioned embedding vectors associated with each cluster further includes: associating, for each user query in the cluster including the user behavior data, the user query to one of a plurality of electronic documents of a third index according to the embedding vectors; and associating, for each user query in the cluster not including the user behavior data, the user query to one of the plurality of electronic documents of the third index based on a frequency of the associations of the user queries including the user behavior data. In some embodiments, the third index includes the plurality of electronic documents of the second index. In some embodiments, the second index and the third index are generated based on the associations.

[0139] In some embodiments, the operations further including: using the machine learning model and determine embedding vectors representative of the plurality of entities; and using the machine learning model and validate the mapping between the cluster and the plurality of entities. In some embodiments, the validation is based on the embedding vectors representative of the user queries in the cluster and the embedding vectors representative of the plurality of entities.

[0140] In some embodiments, the first log data and the second log data further includes at least one of: account data, application data, location data, or language preference data.

[0141] In some embodiments, the subset of the plurality of entities of the second index are provided to the given user based on applying at least one of: a fuzzy search algorithm, a prefix search algorithm, or a precise search algorithm using the machine learning model.

[0142] In some embodiments, each of the plurality of entities includes text data including at least one of: an entity name, a location, a unique identifier, or one or more aliases. In some embodiments, the one or more aliases are representative of the entity name based on the user queries by the plurality of users.

[0143] In some embodiments, a computer-readable media having stored thereon instructions that are executable by a processor of a computing device to enable the processor to perform operations including: obtain a text input for user queries from a plurality of users, each of the user queries corresponding to a request for one of a plurality of electronic documents; retrieve a first log data based on the user queries, the first log data including user behavior data associated with the plurality of users; using a machine learning model and determine embedding vectors representative of the user queries based on the text input and the first log data; partition the embedding vectors into a respective one of a plurality of clusters according to characteristics of each of the embedding vectors, each cluster including one or more of the user queries; identify one or more clusters of the plurality of clusters based on the embedding vectors being less than a predefined threshold; refine a search space by filtering the identified one or more clusters from the plurality of clusters; map each of the remaining plurality of clusters to a respective one of the plurality of electronic documents according to the partitioned embedding vectors associated with each cluster; generate a first index representative of a frequency each of the plurality of electronic documents is queried by the plurality of users according to the partitioned clusters and a second index representative of the plurality of electronic documents according to the mapping; and store the first index and the second index for fulfilling subsequent user queries corresponding to requests for return of at least one of the plurality of electronic documents from users. In some embodiments, for the remaining plurality of clusters, a user query in a cluster having embedding vectors representative of the user behavior data is mapped to an electronic document in a third index based on embedding vectors values. In some embodiments, for the remaining plurality of clusters, a user query in the cluster not having embedding vectors representative of the user behavior data is mapped to an electronic document in the third index based on the frequency each of the plurality of electronic documents is queried by the plurality of users according to the first index. In some embodiments, the second index and the third index are generated based on the mapping of the remaining plurality of clusters.

[0144] In some embodiments, the operations further including: using the machine learning model and determine embedding vectors representative of the at least one of the plurality of electronic documents; and using the machine learning model and validate the mapping between the cluster and the at least one of the plurality of electronic documents. In some embodiments, the validation is based on the embedding vectors values representative of the user queries in the cluster and the embedding vectors values representative of the plurality of electronic documents. In some embodiments, each of the plurality of electronics documents includes text data including at least one of:an entity name, a location, a unique identifier, or one or more aliases. In some embodiments, the one or more aliases are representative of the entity name based on the user queries by the plurality of users.

[0145] In some embodiments, the operations further including: receiving a text input of a given user query from a user computing device of a given user; retrieving, in response to the given user query, a second log data including user behavior data associated with the given user; using the machine learning model and ranking the plurality of electronic documents of the second index based on the first index and the second log data; using the machine learning model and providing a subset of the ranked plurality of electronic documents from the second index to the given user; and storing the subset of the ranked plurality of electronic documents and user interaction data received from the given user in response to the subset of the ranked plurality of electronic documents as user behavior data for subsequent user queries. In some embodiments, the subset of the ranked plurality of electronic documents from the second index are provided to the given user based on the machine learning model using at least one of: a fuzzy search algorithm, a prefix search algorithm, or a precise search algorithm. In some embodiments, the first log data and the second log data further includes at least one of: account data; application data; location data; or language preference data.

[0146] All prior patents and publications referenced herein are incorporated by reference in their entireties.

[0147] Throughout the specification and claims, the following terms take the meanings explicitly associated herein, unless the context clearly dictates otherwise. The phrases "in one embodiment, " “in an embodiment, ” and "in some embodiments" as used herein do not necessarily refer to the same embodiment (s) , though it may. Furthermore, the phrases "in another embodiment" and "in some other embodiments" as used herein do not necessarily refer to a different embodiment, although it may. All embodiments of the disclosure are intended to be combinable without departing from the scope or spirit of the disclosure.

[0148] Some portions of the detailed descriptions of this disclosure have been presented in terms of procedures, logic blocks, processing, and other symbolic representations of operations on data bits within a computer or digital system memory. These descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. A procedure, logic block, process, etc., is herein, and generally, conceived to be a self-consistent sequence of steps or instructions leading to a desired result. The steps are those requiring physical manipulations of physical quantities. Usually, though not necessarily, these physical manipulations take the form of electrical or magnetic data capable of being stored, transferred, combined, compared, and otherwise manipulated in a computer system or similar electronic computing device. For reasons of convenience, and with reference to common usage, such data is referred to as bits, values, elements, symbols, characters, terms, numbers, or the like, with reference to various presently disclosed embodiments. It should be borne in mind, however, that these terms are to be interpreted as referencing physical manipulations and quantities and are merely convenient labels that should be interpreted further in view of terms commonly used in the art. Unless specifically stated otherwise, as apparent from the discussion herein, it is understood that throughout discussions of the present embodiment, discussions utilizing terms such as “determining” or “outputting” or “transmitting” or “recording” or “locating” or “storing” or “displaying” or “receiving” or “recognizing” or “utilizing” or “generating” or “providing” or “accessing” or “checking” or “notifying” or “delivering” or the like, refer to the action and processes of a computer system, or similar electronic computing device, that manipulates and transforms data. The data is represented as physical (electronic) quantities within the computer system’s registers and memories and is transformed into other data similarly represented as physical quantities within the computer system memories or registers, or other such information storage, transmission, or display devices as described herein or otherwise understood to one of ordinary skill in the art.

[0149] As used herein, the term "based on" is not exclusive and allows for being based on additional factors not described, unless the context clearly dictates otherwise. In addition, throughout the specification, the meaning of "a, " "an, " and "the" include plural references. The meaning of "in" includes "in" and "on. "

[0150] ASPECTS

[0151] Various Aspects are described below. It is to be understood that any one or more of the features recited in the following Aspect (s) can be combined with any one or more other Aspect (s) .

[0152] Aspect 1. A computer-implemented method comprising: obtaining a text input for user queries from a plurality of users, each of the user queries corresponding to a request for one of a plurality of electronic documents; retrieving a first log data based on the user queries, the first log data comprising user behavior data associated with the plurality of users; using a machine learning model and determining embedding vectors representative of the user queries based on the text input and the first log data; partitioning the embedding vectors into a respective one of a plurality of clusters according to characteristics of each of the embedding vectors, each cluster including one or more of the user queries; mapping each of the plurality of clusters to a respective one of the plurality of electronic documents according to the partitioned embedding vectors associated with each cluster; generating a first index representative of a frequency each of the plurality of electronic documents is queried by the plurality of users according to the partitioned embedding vectors and a second index representative of the plurality of electronic documents according to the mapping; and storing the first index and the second index for fulfilling subsequent user queries corresponding to requests for return of at least one of the plurality of electronic documents from users.

[0153] Aspect 2. The computer-implemented method according to aspect 1, further comprising: using the machine learning model and determining embedding vectors representative of the plurality of electronic documents; and using the machine learning model and validating the mapping between the cluster and the plurality of electronic documents, wherein the validation is based on the embedding vectors representative of the user queries in the cluster and the embedding vectors representative of the plurality of electronic documents.

[0154] Aspect 3. The computer-implemented method according to any of the preceding aspects, further comprising: identifying one or more clusters of the plurality of clusters based on the embedding vectors being less than a predefined threshold; and refining a search space by filtering the identified one or more clusters from the plurality of clusters, wherein the mapping comprises mapping each of the remaining clusters of the plurality of clusters to a respective one of the plurality of electronic documents.

[0155] Aspect 4. The computer-implemented method according to any of the preceding aspects, wherein mapping each of the plurality of clusters to the respective one of the plurality of electronic documents according to the partitioned embedding vectors associated with each cluster further comprises: associating, for each user query in the cluster including the user behavior data, the user query to one of a plurality of electronic documents of a third index according to the embedding vectors; and associating, for each user query in the cluster not including the user behavior data, the user query to one of the plurality of electronic documents of the third index based on a frequency of the associations of the user queries including the user behavior data, wherein the third index includes the plurality of electronic documents of the second index, and wherein the second index and the third index are generated based on the associations.

[0156] Aspect 5. The computer-implemented method according to any of the preceding aspects, further comprising: receiving a text input of a given user query from a user computing device of a given user; retrieving, in response to the given user query, a second log data comprising user behavior data associated with the given user; using the machine learning model and ranking the plurality of electronic documents of the second index based on the first index and the second log data; using the machine learning model and providing a subset of the ranked plurality of electronic documents from the second index to the given user; and storing the subset of the ranked plurality of electronic documents for fulfilling the subsequent user queries.

[0157] Aspect 6. The computer-implemented method according to aspect 5, wherein the second log data comprises data associated with the given user comprising at least one of:account data; application data; location data; or language preference data.

[0158] Aspect 7. The computer-implemented method according to aspect 5, wherein the subset of the plurality of electronic documents of the second index are provided to the given user based on the machine learning model using at least one of: a fuzzy search algorithm, a prefix search algorithm, or a precise search algorithm.

[0159] Aspect 8. The computer-implemented method according to any of the preceding aspects, wherein the embedding vectors are representative of semantic characteristics of the text input and representative of contextual relationships of the first log data.

[0160] Aspect 9. The computer-implemented method according to any of the preceding aspects, wherein the first log data comprises data associated with the plurality of users comprising at least one of: account data; application data; location data; or language preference data.

[0161] Aspect 10. The computer-implemented method according to any of the preceding aspects, wherein each of the plurality of electronics documents comprises text data comprising at least one of: an entity name; a location; a unique identifier; or one or more aliases, the one or more aliases are representative of the entity name based on the user queries by the plurality of users.

[0162] Aspect 11. A system comprising: a processor; and a non-transitory computer readable medium having stored thereon instructions that are executable by the processor to perform operations comprising: obtain a text input for user queries from a plurality of users, each of the user queries corresponding to a request for one of a plurality of entities; retrieve a first log data based on the user queries, the first log data comprising user behavior data associated with the plurality of users; use a machine learning model and generate a first index representative of a frequency each of the plurality of entities is queried by the plurality of users according to a clustering of the user queries and a second index representative of a mapping of the clusters of the user queries to the plurality of entities; and store the first index and the second index for fulfilling subsequent user queries corresponding to requests for return of at least one of the plurality of entities from users, wherein the subsequent user queries comprises: receive a text input of a given user query from a user computing device of a given user; retrieve, in response to the given user query, a second log data comprising user behavior data associated with the given user; using the machine learning model and rank the plurality of entities of the second index based on the first index and the second log data; and using the machine learning model and provide a subset of the ranked plurality of entities from the second index to the given user.

[0163] Aspect 12. The system according to aspect 11, wherein the operations further comprising: using a machine learning model and determine embedding vectors representative of the user queries based on the text input and the first log data; partition the embedding vectors into a respective one of a plurality of clusters according to characteristics of each of the embedding vectors, each cluster including one or more of the user queries; identify one or more clusters of the plurality of clusters based on the embedding vectors representative of the user behavior data in the one or more clusters being less than a predefined threshold; refine a search space by filtering the identified one or more clusters from the plurality of clusters; and mapping each of the remaining plurality of clusters to a respective one of the plurality of entities according to the partitioned embedding vectors associated with each cluster.

[0164] Aspect 13. The system according to aspects 11 or 12, wherein mapping each of the remaining plurality of clusters to the respective one of the plurality of entities according to the partitioned embedding vectors associated with each cluster further comprises: associating, for each user query in the cluster including the user behavior data, the user query to one of a plurality of electronic documents of a third index according to the embedding vectors; and associating, for each user query in the cluster not including the user behavior data, the user query to one of the plurality of electronic documents of the third index based on a frequency of the associations of the user queries including the user behavior data, wherein the third index includes the plurality of electronic documents of the second index, and wherein the second index and the third index are generated based on the associations.

[0165] Aspect 14. The system according to aspects 11, 12, or 13, wherein the operations further comprising: using the machine learning model and determine embedding vectors representative of the plurality of entities; and using the machine learning model and validate the mapping between the cluster and the plurality of entities, wherein the validation is based on the embedding vectors representative of the user queries in the cluster and the embedding vectors representative of the plurality of entities.

[0166] Aspect 15. The system according to aspects 11, 12, 13, or 14, wherein the first log data and the second log data further comprises at least one of: account data; application data; location data; or language preference data.

[0167] Aspect 16. The system according to aspects 11, 12, 13, 14, or 15, wherein the subset of the plurality of entities of the second index are provided to the given user based on applying at least one of a fuzzy search algorithm, a prefix search algorithm, or a precise search algorithm using the machine learning model.

[0168] Aspect 17. The system according to aspects 11, 12, 13, 14, 15, or 16, wherein each of the plurality of entities comprises text data comprising at least one of: an entity name; a location; a unique identifier; or one or more aliases, the one or more aliases are representative of the entity name based on the user queries by the plurality of users.

[0169] Aspect 18. A computer-readable media having stored thereon instructions that are executable by a processor of a computing device to enable the processor to perform operations comprising: obtain a text input for user queries from a plurality of users, each of the user queries corresponding to a request for one of a plurality of electronic documents; retrieve a first log data based on the user queries, the first log data comprising user behavior data associated with the plurality of users; using a machine learning model and determine embedding vectors representative of the user queries based on the text input and the first log data; partition the embedding vectors into a respective one of a plurality of clusters according to characteristics of each of the embedding vectors, each cluster including one or more of the user queries; identify one or more clusters of the plurality of clusters based on the embedding vectors being less than a predefined threshold; refine a search space by filtering the identified one or more clusters from the plurality of clusters; map each of the remaining plurality of clusters to a respective one of the plurality of electronic documents according to the partitioned embedding vectors associated with each cluster; generate a first index representative of a frequency each of the plurality of electronic documents is queried by the plurality of users according to the partitioned clusters and a second index representative of the plurality of electronic documents according to the mapping; and store the first index and the second index for fulfilling subsequent user queries corresponding to requests for return of at least one of the plurality of electronic documents from users, wherein, for the remaining plurality of clusters, a user query in a cluster having embedding vectors representative of the user behavior data is mapped to an electronic document in a third index based on embedding vectors values, wherein, for the remaining plurality of clusters, a user query in the cluster not having embedding vectors representative of the user behavior data is mapped to an electronic document in the third index based on the frequency each of the plurality of electronic documents is queried by the plurality of users according to the first index, and wherein the second index and the third index are generated based on the mapping of the remaining plurality of clusters.

[0170] Aspect 19. The computer-readable media according to aspect 18, the operations further comprising: using the machine learning model and determine embedding vectors representative of the at least one of the plurality of electronic documents; and using the machine learning model and validate the mapping between the cluster and the at least one of the plurality of electronic documents, wherein the validation is based on the embedding vectors values representative of the user queries in the cluster and the embedding vectors values representative of the plurality of electronic documents, and wherein each of the plurality of electronics documents comprises text data comprising at least one of: an entity name; a location; a unique identifier; or one or more aliases, the one or more aliases are representative of the entity name based on the user queries by the plurality of users.

[0171] Aspect 20. The computer-readable media according to aspects 18 or 19, the operations further comprising: receiving a text input of a given user query from a user computing device of a given user; retrieving, in response to the given user query, a second log data comprising user behavior data associated with the given user; using the machine learning model and ranking the plurality of electronic documents of the second index based on the first index and the second log data; using the machine learning model and providing a subset of the ranked plurality of electronic documents from the second index to the given user; and storing the subset of the ranked plurality of electronic documents and user interaction data received from the given user in response to the subset of the ranked plurality of electronic documents as user behavior data for subsequent user queries, wherein the subset of the ranked plurality of electronic documents from the second index are provided to the given user based on the machine learning model using at least one of: a fuzzy search algorithm, a prefix search algorithm, or a precise search algorithm, wherein the first log data and the second log data further comprises at least one of: account data; application data; location data;or language preference data.

Claims

1.A computer-implemented method comprising:obtaining a text input for user queries from a plurality of users, each of the user queries corresponding to a request for one of a plurality of electronic documents;retrieving a first log data based on the user queries, the first log data comprising user behavior data associated with the plurality of users;using a machine learning model and determining embedding vectors representative of the user queries based on the text input and the first log data;partitioning the embedding vectors into a respective one of a plurality of clusters according to characteristics of each of the embedding vectors, each cluster including one or more of the user queries;mapping each of the plurality of clusters to a respective one of the plurality of electronic documents according to the partitioned embedding vectors associated with each cluster;generating a first index representative of a frequency each of the plurality of electronic documents is queried by the plurality of users according to the partitioned embedding vectors and a second index representative of the plurality of electronic documents according to the mapping; andstoring the first index and the second index for fulfilling subsequent user queries corresponding to requests for return of at least one of the plurality of electronic documents from users.2.The computer-implemented method of claim 1, further comprising:using the machine learning model and determining embedding vectors representative of the plurality of electronic documents; andusing the machine learning model and validating the mapping between the cluster and the plurality of electronic documents,wherein the validation is based on the embedding vectors representative of the user queries in the cluster and the embedding vectors representative of the plurality of electronic documents.3.The computer-implemented method of claim 1, further comprising:identifying one or more clusters of the plurality of clusters based on the embedding vectors being less than a predefined threshold; andrefining a search space by filtering the identified one or more clusters from the plurality of clusters,wherein the mapping comprises mapping each of the remaining clusters of the plurality of clusters to a respective one of the plurality of electronic documents.4.The computer-implemented method of claim 3, wherein mapping each of the plurality of clusters to the respective one of the plurality of electronic documents according to the partitioned embedding vectors associated with each cluster further comprises:associating, for each user query in the cluster including the user behavior data, the user query to one of a plurality of electronic documents of a third index according to the embedding vectors; andassociating, for each user query in the cluster not including the user behavior data, the user query to one of the plurality of electronic documents of the third index based on a frequency of the associations of the user queries including the user behavior data,wherein the third index includes the plurality of electronic documents of the second index, and wherein the second index and the third index are generated based on the associations.5.The computer-implemented method of claim 1, further comprising:receiving a text input of a given user query from a user computing device of a given user;retrieving, in response to the given user query, a second log data comprising user behavior data associated with the given user;using the machine learning model and ranking the plurality of electronic documents of the second index based on the first index and the second log data;using the machine learning model and providing a subset of the ranked plurality of electronic documents from the second index to the given user; andstoring the subset of the ranked plurality of electronic documents for fulfilling the subsequent user queries.6.The computer-implemented method of claim 5, wherein the second log data comprises data associated with the given user comprising at least one of:account data;application data;location data; orlanguage preference data.7.The computer-implemented method of claim 5, wherein the subset of the plurality of electronic documents of the second index are provided to the given user based on the machine learning model using at least one of: a fuzzy search algorithm, a prefix search algorithm, or a precise search algorithm.8.The computer-implemented method of claim 1, wherein the embedding vectors are representative of semantic characteristics of the text input and representative of contextual relationships of the first log data.9.The computer-implemented method of claim 1, wherein the first log data comprises data associated with the plurality of users comprising at least one of:account data;application data;location data; orlanguage preference data.10.The computer-implemented method of claim 1, wherein each of the plurality of electronics documents comprises text data comprising at least one of:an entity name;a location;a unique identifier; orone or more aliases, the one or more aliases are representative of the entity name based on the user queries by the plurality of users.11.A system comprising:a processor; anda non-transitory computer readable medium having stored thereon instructions that are executable by the processor to perform operations comprising:obtain a text input for user queries from a plurality of users, each of the user queries corresponding to a request for one of a plurality of entities;retrieve a first log data based on the user queries, the first log data comprising user behavior data associated with the plurality of users;use a machine learning model and generate a first index representative of a frequency each of the plurality of entities is queried by the plurality of users according to a clustering of the user queries and a second index representative of a mapping of the clusters of the user queries to the plurality of entities; andstore the first index and the second index for fulfilling subsequent user queries corresponding to requests for return of at least one of the plurality of entities from users, wherein the subsequent user queries comprises:receive a text input of a given user query from a user computing device of a given user;retrieve, in response to the given user query, a second log data comprising user behavior data associated with the given user;using the machine learning model and rank the plurality of entities of the second index based on the first index and the second log data; andusing the machine learning model and provide a subset of the ranked plurality of entities from the second index to the given user.12.The system of claim 11, wherein the operations further comprising:using a machine learning model and determine embedding vectors representative of the user queries based on the text input and the first log data;partition the embedding vectors into a respective one of a plurality of clusters according to characteristics of each of the embedding vectors, each cluster including one or more of the user queries;identify one or more clusters of the plurality of clusters based on the embedding vectors representative of the user behavior data in the one or more clusters being less than a predefined threshold;refine a search space by filtering the identified one or more clusters from the plurality of clusters; andmapping each of the remaining plurality of clusters to a respective one of the plurality of entities according to the partitioned embedding vectors associated with each cluster.13.The system of claim 12, wherein mapping each of the remaining plurality of clusters to the respective one of the plurality of entities according to the partitioned embedding vectors associated with each cluster further comprises:associating, for each user query in the cluster including the user behavior data, the user query to one of a plurality of electronic documents of a third index according to the embedding vectors; andassociating, for each user query in the cluster not including the user behavior data, the user query to one of the plurality of electronic documents of the third index based on a frequency of the associations of the user queries including the user behavior data,wherein the third index includes the plurality of electronic documents of the second index, and wherein the second index and the third index are generated based on the associations.14.The system of claim 11, wherein the operations further comprising:using the machine learning model and determine embedding vectors representative of the plurality of entities; andusing the machine learning model and validate the mapping between the cluster and the plurality of entities,wherein the validation is based on the embedding vectors representative of the user queries in the cluster and the embedding vectors representative of the plurality of entities.15.The system of claim 11, wherein the first log data and the second log data further comprises at least one of:account data;application data;location data; orlanguage preference data.16.The system of claim 11, wherein the subset of the plurality of entities of the second index are provided to the given user based on applying at least one of a fuzzy search algorithm, a prefix search algorithm, or a precise search algorithm using the machine learning model.17.The system of claim 11, wherein each of the plurality of entities comprises text data comprising at least one of:an entity name;a location;a unique identifier; orone or more aliases, the one or more aliases are representative of the entity name based on the user queries by the plurality of users.18.A computer-readable media having stored thereon instructions that are executable by a processor of a computing device to enable the processor to perform operations comprising:obtain a text input for user queries from a plurality of users, each of the user queries corresponding to a request for one of a plurality of electronic documents;retrieve a first log data based on the user queries, the first log data comprising user behavior data associated with the plurality of users;using a machine learning model and determine embedding vectors representative of the user queries based on the text input and the first log data;partition the embedding vectors into a respective one of a plurality of clusters according to characteristics of each of the embedding vectors, each cluster including one or more of the user queries;identify one or more clusters of the plurality of clusters based on the embedding vectors being less than a predefined threshold;refine a search space by filtering the identified one or more clusters from the plurality of clusters;map each of the remaining plurality of clusters to a respective one of the plurality of electronic documents according to the partitioned embedding vectors associated with each cluster;generate a first index representative of a frequency each of the plurality of electronic documents is queried by the plurality of users according to the partitioned clusters and a second index representative of the plurality of electronic documents according to the mapping; andstore the first index and the second index for fulfilling subsequent user queries corresponding to requests for return of at least one of the plurality of electronic documents from users,wherein, for the remaining plurality of clusters, a user query in a cluster having embedding vectors representative of the user behavior data is mapped to an electronic document in a third index based on embedding vectors values, wherein, for the remaining plurality of clusters, a user query in the cluster not having embedding vectors representative of the user behavior data is mapped to an electronic document in the third index based on the frequency each of the plurality of electronic documents is queried by the plurality of users according to the first index, and wherein the second index and the third index are generated based on the mapping of the remaining plurality of clusters.19.The computer-readable media of claim 18, the operations further comprising:using the machine learning model and determine embedding vectors representative of the at least one of the plurality of electronic documents; andusing the machine learning model and validate the mapping between the cluster and the at least one of the plurality of electronic documents,wherein the validation is based on the embedding vectors values representative of the user queries in the cluster and the embedding vectors values representative of the plurality of electronic documents, andwherein each of the plurality of electronics documents comprises text data comprising at least one of:an entity name;a location;a unique identifier; orone or more aliases, the one or more aliases are representative of the entity name based on the user queries by the plurality of users.20.The computer-readable media of claim 18, the operations further comprising:receiving a text input of a given user query from a user computing device of a given user;retrieving, in response to the given user query, a second log data comprising user behavior data associated with the given user;using the machine learning model and ranking the plurality of electronic documents of the second index based on the first index and the second log data;using the machine learning model and providing a subset of the ranked plurality of electronic documents from the second index to the given user; andstoring the subset of the ranked plurality of electronic documents and user interaction data received from the given user in response to the subset of the ranked plurality of electronic documents as user behavior data for subsequent user queries,wherein the subset of the ranked plurality of electronic documents from the second index are provided to the given user based on the machine learning model using at least one of: a fuzzy search algorithm, a prefix search algorithm, or a precise search algorithm,wherein the first log data and the second log data further comprises at least one of: account data; application data; location data; or language preference data.