Evaluating the Interpretation of Search Queries
By using the training model to evaluate the manual interpretation of search queries using the training data set, the problem of evaluating the interpretation quality of natural language search queries in the prior art is solved, and efficient and accurate evaluation is achieved, which is suitable for large-scale locale environments.
Patent Information
- Application Number
- CN202080103390.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-01
- Filing Date
- 2020-12-10
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-12-10
AI Technical Summary
The prior art has difficulty in effectively evaluating the quality of interpretation of natural language search queries, especially with challenges in scale and cost-effectiveness.
Evaluate whether the manual interpretation of the search query is correct by training the model to leverage the training dataset, including past search queries, manual interpretation, and evaluation tags. The method includes receiving search queries and manual interpretation, performing initial evaluation using the initial model, and generating a final evaluation in combination with temporal features and clustering features.
It improves the efficiency and accuracy of the quality of natural language search query interpretation, reduces the cost and time of manual evaluation, and is suitable for large-scale locale environments.
Smart Images

Figure CN116113959B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of priority to U.S. Provisional Application Serial No. 63 / 047,039, filed Jul. 1, 2020, the content of which is incorporated herein by reference. Technical Field
[0003] This disclosure relates to evaluating interpretations of search queries. Background Art
[0004] In many systems, measuring the quality of an interpretation of a search query involves using user interface interactions (such as the selection of results for a search query) as a proxy for correctly interpreting the search query and providing appropriate results. However, in some implementations of answering natural language search queries, once a search query is input, the results of the search query are designed in a panel. As such, interaction signals cannot be analyzed to gauge the quality of the results provided in response to the search query. Another method of measuring the quality of an answer to a search query is to use human evaluation to rate the quality of the answer. However, sending the interpretation and results of a search query to a human to manually rate the quality can be expensive and time - consuming. Additionally, when search queries are scaled for a large number of languages, it becomes extremely difficult to use humans to evaluate the quality of interpretations of search queries. Accordingly, there is a desire for a predictive method in terms of evaluating the accuracy of interpretations of natural language search queries to increase efficiency, improve the cost of evaluating analysis data quality, and improve the quality of the results returned by search queries. Summary of the Invention
[0005] One aspect of the present disclosure relates to a method for evaluating the accuracy of a human interpretation of a search query. The method may include receiving, by one or more processing circuits, a training data set. The training data set may include a plurality of past search queries, a human interpretation of each of the plurality of past search queries, and a human evaluation label indicating whether the human interpretation of each of the plurality of past search queries is correct. The method may include training, by one or more processing circuits, a first model using the training data set. The first model may be trained to evaluate whether a human interpretation of a search query is correct. The method may include receiving, by one or more processing circuits, a search query and a human interpretation of the search query, the search query including one or more words entered via a user interface to request desired information. The human interpretation may include one or more words defining an interpretation of the desired information. The method may include determining, by one or more processing circuits, an initial evaluation of whether the human interpretation of the search query is correct using the first model. The method may include generating, by one or more processing circuits, a second model using the initial evaluation from the first model, time features associated with the search query, and clustering features associated with the search query. The method may include determining, by one or more processing circuits, a final evaluation of whether the human interpretation of the search query is correct using the second model.
[0006] In some embodiments of the method, the search query may be a first search query. In some embodiments of the method, the method may include determining, by one or more processing circuits, whether the first search query is associated with a second search query received within a predetermined time interval after the first search query. In some embodiments of the method, the method may include receiving, by one or more processing circuits, a token embedding from the first model for each of the first search query and the second search query. In some embodiments of the method, the token may be a word in the search query. In some embodiments of the method, the method may include determining, by one or more processing circuits, a vector sentence representation for each of the first search query and the second search query by averaging the token embeddings from the first model for each of the first search query and the second search query.
[0007] In some embodiments of the method, the first model may be pre-trained on a natural language data set and the first model may be trained using the training data set to adapt the first model to a specific classification problem.
[0008] In some embodiments of the method, the method may include parsing, by one or more processing circuits, the first search query and the second search query using a distance algorithm. In some embodiments of the method, the distance algorithm may be at least one of an Euclidean distance algorithm or a cosine similarity algorithm.
[0009] In some embodiments of the method, the method can include determining, by one or more processing circuits, whether a second search query is a search refinement of a first search query. In some embodiments of the method, a search refinement can be a weighted indication of an incorrect human interpretation of a search query.
[0010] In some embodiments of the method, the method can include generating, by one or more processing circuits, a cluster at least partially based on similarities between search queries in a cluster of search queries. In some embodiments of the method, the method can include updating, by one or more processing circuits, the generated cluster in response to receiving a new search query.
[0011] In some embodiments of the method, the method can include determining, by one or more processing circuits, whether an input for viewing a report for a search query is received via a user interface. In some embodiments of the method, the input for viewing a report can be a weighted indication of a correct human interpretation of a search query.
[0012] Another aspect of the present disclosure relates to a system configured to evaluate the accuracy of a human interpretation of a search query. The system can include one or more hardware processors configured by machine-readable instructions. The processor can be configured to receive, by one or more processing circuits, a training data set. The training data set can include a plurality of past search queries, a human interpretation of each of the plurality of past search queries, and a human evaluation label indicating whether the human interpretation of each of the plurality of past search queries is correct. The processor can be configured to train, by one or more processing circuits, a first model using the training data set. The first model can be trained to evaluate whether a human interpretation of a search query is correct. The processor can be configured to receive, by one or more processing circuits, a search query and a human interpretation of the search query, the search query including one or more words input via a user interface to request desired information. The human interpretation can include one or more words defining an interpretation of the desired information. The processor can be configured to determine, by one or more processing circuits, an initial evaluation of whether the human interpretation of the search query is correct using the first model. The processor can be configured to generate, by one or more processing circuits, a second model using the initial evaluation from the first model, time features associated with the search query, and cluster features associated with the search query. The processor can be configured to determine, by one or more processing circuits, a final evaluation of whether the human interpretation of the search query is correct using the second model.
[0013] In some embodiments of the system, the search query can be a first search query. In some embodiments of the system, the processor can be configured to determine, by one or more processing circuits, whether the first search query is associated with a second search query received within a predetermined time interval. In some embodiments of the system, the processor can be configured to receive, by one or more processing circuits, token embeddings from a first model for each of the first search query and the second search query. In some embodiments of the system, a token can be a word in the search query. In some embodiments of the system, the processor can be configured to determine, by one or more processing circuits, a vector sentence representation for each of the first search query and the second search query by averaging the token embeddings from the first model for each of the first search query and the second search query.
[0014] In some embodiments of the system, the first model can be pre-trained on a natural language dataset and tuned to a specific classification problem using a training dataset.
[0015] In some embodiments of the system, the processor can be configured to parse the first search query and the second search query using a distance algorithm by one or more processing circuits. In some embodiments of the system, the distance algorithm can be at least one of an Euclidean distance algorithm or a cosine similarity algorithm.
[0016] In some embodiments of the system, the processor can be configured to determine, by one or more processing circuits, whether the second search query is a search refinement of the first search query. In some embodiments of the system, a search refinement is a weighted indication of an incorrect human interpretation of the search query.
[0017] In some embodiments of the system, the processor can be configured to generate clusters by one or more processing circuits based at least in part on similarities between search queries in a cluster of search queries. In some embodiments of the system, the processor can be configured to update the generated clusters in response to receiving a new search query by one or more processing circuits.
[0018] In some embodiments of the system, the processor can be configured to determine, by one or more processing circuits, whether an input for viewing a report of the search query is received via a user interface. In some embodiments of the system, the input for viewing the report can be a weighted indication of a correct human interpretation of the search query.
[0019] Another aspect of the present disclosure relates to a non - transitory computer - readable storage medium having instructions thereon that are executable by one or more processors to perform operations for evaluating the accuracy of a human interpretation of a search query. The operations can include receiving, by one or more processing circuits, a training data set. The training data set can include a plurality of past search queries, a human interpretation for each of the plurality of past search queries, and a human - evaluation label indicating whether the human interpretation for each of the plurality of past search queries is correct. The operations can include training, by one or more processing circuits, a first model using the training data set. The first model can be trained to evaluate whether a human interpretation of a search query is correct. The operations can include receiving, by one or more processing circuits, a search query and a human interpretation of the search query, the search query including one or more words entered via a user interface to request desired information. The human interpretation can include one or more words defining an interpretation of the desired information. The operations can include determining, by one or more processing circuits, an initial evaluation of whether the human interpretation of the search query is correct using the first model. The operations can include generating, by one or more processing circuits, a second model using the initial evaluation from the first model, temporal features associated with the search query, and clustering features associated with the search query. The operations can include determining, by one or more processing circuits, a final evaluation of whether the human interpretation of the search query is correct using the second model.
[0020] In some embodiments of the computer - readable storage medium, the search query can be a first search query. In some embodiments of the computer - readable storage medium, the operations can include determining, by one or more processing circuits, whether the first search query is associated with a second search query received within a predetermined time interval. In some embodiments of the computer - readable storage medium, the operations can include receiving, by one or more processing circuits, token embeddings from the first model for each of the first search query and the second search query. In some embodiments of the computer - readable storage medium, a token can be a word in the search query. In some embodiments of the computer - readable storage medium, the operations can include determining, by one or more processing circuits, a vector sentence representation for each of the first search query and the second search query by averaging the token embeddings from the first model for each of the first search query and the second search query.
[0021] In some embodiments of the computer - readable storage medium, the first model can be pre - trained on a natural - language data set, and training the first model using the training data set tunes the first model to a specific classification problem.
[0022] In some embodiments of the computer-readable storage medium, the operation may include parsing a first search query and a second search query by one or more processing circuits using a distance algorithm. In some embodiments of the computer-readable storage medium, the distance algorithm may be at least one of an Euclidean distance algorithm or a cosine similarity algorithm.
[0023] In some embodiments of the computer-readable storage medium, the operation may include determining by one or more processing circuits whether the second search query is a search refinement of the first search query. In some embodiments of the computer-readable storage medium, a search refinement may be a weighted indication of an incorrect human interpretation of a search query.
[0024] In some embodiments of the computer-readable storage medium, the operation may include generating a cluster by one or more processing circuits based at least in part on similarities between search queries in a cluster of search queries. In some embodiments of the computer-readable storage medium, the operation may include updating the generated cluster by one or more processing circuits in response to receiving a new search query. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
[0026] Figure 1 is a block diagram of a system configured to evaluate the accuracy of a human interpretation of a search query and an associated environment, according to an illustrative embodiment.
[0027] Figure 2 is according to an illustrative embodiment Figure 1 of a system configured to evaluate the accuracy of a human interpretation of a search query.
[0028] Figure 3 is a flowchart of a process for evaluating the accuracy of a human interpretation of a search query, according to an illustrative embodiment.
[0029] Figure 4 is a flowchart of a process for analyzing a search query to evaluate whether a human interpretation of the search query is an accurate interpretation, according to an illustrative embodiment.
[0030] Figure 5 is a flowchart of a process for evaluating the accuracy of a human interpretation of a search query, according to an illustrative embodiment.
[0031] Figure 6 is a flowchart of a process for evaluating the accuracy of a human interpretation of a search query using clustering, according to an illustrative embodiment.
[0032] Figure 7 is a block diagram of the inputs and outputs of a first model and a final model utilized by a system for Figure 1 .
[0033] Figure 8 is a diagram of the structure of a first model utilized by a system for Figure 1 in accordance with an illustrative embodiment.
[0034] Figure 9 is a user interface for entering a search query and viewing a human interpretation and results of the search query in accordance with an illustrative embodiment.
[0035] Figure 10 is a block diagram of a computing system in accordance with an illustrative embodiment. DETAILED DESCRIPTION
[0036] The following is a more detailed description of various concepts related to evaluating the quality of human interpretations of natural language search queries and embodiments of methods, apparatuses, and systems for providing information using a computer network. Since the concepts described are not limited to any particular embodiment, the various concepts introduced above and discussed in more detail below can be implemented in any of a variety of ways. Examples of specific embodiments and applications are provided primarily for illustrative purposes.
[0037] Generally referring to the drawings, various illustrative systems and methods for providing query results to content providers are shown. More specifically, the present disclosure relates to systems and methods for evaluating the accuracy of human interpretations of search queries (e.g., natural language search queries).
[0038] To produce quality data results for a search query, the systems and methods depend on a correct interpretation of the type of information that a user expects from the search query. Thus, the systems and methods described herein relate to a method for automatically evaluating whether a human interpretation of a natural language search query is a correct interpretation of the expected type of information. For example, the search query "What pages do people spend the most time on" has the correct human interpretation "The pages with the most average time on page". As another example, the search query "Slowest page" may have an incorrect human interpretation "The page with the fewest page views" and a correct human interpretation "The page with the most average page load time".
[0039] In a typical system, user interface interactions (such as selecting (e.g., clicking) a particular result) are analyzed to determine whether a search query has been correctly interpreted, so as to provide more relevant and accurate results. However, in some systems, user interaction signals cannot be considered to determine the quality of the interpretation of a search query. For example, in some systems, once a user enters a search query, a single best result for the search query can be selected and displayed in a result panel. In other systems, manual input from humans is used to evaluate the quality of the answer to a search query. Using individuals to rate the accuracy of the interpretation of a search query can be time-consuming and costly. In addition, as the volume of search queries increases to include several languages, using manual evaluation to determine the quality of the interpretation of a search query is particularly challenging. Accordingly, a more scalable predictive method for rating interpretations is disclosed to beneficially reduce the time and cost required in evaluating the quality of answers to natural language search queries.
[0040] In addition, it is desirable to answer a search query in a way that prevents the user from needing to enter a subsequent search query because the first search query was incorrectly answered. If the system incorrectly interprets the search query, then another search query needs to be entered and processed to provide new results for the subsequent search query. This additional processing may waste valuable server and network resources, resulting in lower efficiency of the overall computing system, which can be increasingly expensive in response to scaling the problem and solution to multiple languages. Given the frequency of search queries that occur per second, avoiding re-entering a search query can prevent unnecessary network packets from being sent and inefficient use of server time. Therefore, reducing the frequency of re-searching a previously entered search query can advantageously improve network utilization and bandwidth efficiency. Additionally, a method for preventing the user from needing to enter a subsequent search query because the first search query was incorrectly answered can reduce the chance that the user abandons the analysis service due to inaccurate results.
[0041] This prediction method allows a natural language processing (NLP) model (such as a bidirectional encoder representations from transformers (BERT) model) to be pre-trained using a large amount of language data (such as Wikipedia) and then further trained (i.e., fine-tuned) using a dataset of search queries and human-rated interpretations of the search queries. After training the NLP model to evaluate whether the human interpretation of a search query is correct, the prediction method generates a new final model (such as a logistic regression model) that is used to determine the final evaluation (e.g., prediction) of whether the human interpretation of a search query is correct. The NLP model is also used to embed search queries for clustering. These clustering features are then used as inputs to the final model, such as the average frequency of search queries issued in the cluster, how many different users issued search queries in the cluster, and so on. Embeddings from the NLP model are also used to determine whether similar search queries are re-issued within a shorter time period along with other temporal features related to the search query, which can be used as an indication that the original human interpretation of the search query is incorrect.
[0042] An advantage of fine-tuning the NLP model can be to allow a smaller training dataset of search queries and human-rated human interpretations, but still produce an accurate model. The model can still maintain a high level of accurate evaluation because the overall pre-trained model can already effectively represent many NLP concepts. Thus, the prediction method disclosed herein generates an evaluation (e.g., prediction) that is still as accurate as in previous methods, but requires much less manual classification to train the model.
[0043] In this disclosure, the terms "human interpretation", "interpretation", and "answer" of a search query can be used interchangeably to refer to an interpretation of what information is expected from the search query. For example, the search query "slowest page" may have a given interpretation "the page with the fewest page views". As another example, the search query "country with the most users" may have a given answer "the country with the most users". Additionally, the terms "query", "search query", and "question" can be used interchangeably herein to describe a user's input of expected information in a single search.
[0044] For systems discussed herein that collect and / or utilize personal information about users, or for situations where personal information may be utilized, users may be provided with an opportunity to control whether programs or features collect personal information (e.g., information about a user's social network, social actions or activities, a user's preferences, a user's current location, etc.) or to control whether and / or how content that may be more relevant to the user is received from a content server. Additionally, before certain data is stored or used, it may be anonymized in one or more ways so that personally identifiable information is removed when generating parameters (e.g., demographic parameters). For example, a user's identity may be anonymized so that the user's personally identifiable information cannot be determined, or a user's geographic location may be generalized to a location where the location information was obtained (such as at the city, zip code, or state level) so that the user's specific location cannot be determined. Accordingly, a user may control how the content server collects and uses information about him or her. Additionally, personal user information itself is not disclosed to content providers, so content providers cannot discern interactions associated with a particular user.
[0045] Now referring to Figure 1 , generally speaking, a block diagram of a computing environment for evaluating the accuracy of a human interpretation of a search query (e.g., a natural language search query) according to an illustrative embodiment is shown. A user may use one or more user devices 104 to perform various actions and / or access various types of content, some of which may be provided via a network 102 (e.g., the Internet, a LAN, a WAN, etc.). For example, the user device 104 may be used to access web pages (e.g., using an Internet browser), media files, and / or any other type of content. A content management system 108 may be configured to select content to display to a user within a resource (e.g., a web page, an application, etc.) and provide content items 128 from a content database 126 to the user device 104 for display within the resource. The content items 128 from which the content management system 108 selects may be provided by one or more content providers using one or more content provider devices 106 via the network 102. In some embodiments, the content management system 108 may select one or more content items 128 from one or more content providers among multiple content items from multiple content providers. In such embodiments, the content management system 108 may determine the content to be published in one or more content interfaces of a resource (e.g., a web page, an application, etc.) to be displayed on the user device 104 based at least in part on metrics or other characteristics of the content item or content provider.
[0046] Referring more specifically to Figure 1, the user device 104 and / or the content provider device 106 can be any type of computing device (e.g., having a processor and a memory or other types of computer-readable storage media), such as a television and / or a set-top box, a mobile communication device (e.g., a cellular phone, a smart phone, etc.), a computer and / or a media device (a desktop computer, a laptop or notebook computer, a netbook computer, a tablet device, a gaming system, etc.), or any other type of computing device. In some embodiments, one or more user devices 104 can be a set-top box or other device used with a television. In some embodiments, content can be provided via web-based applications and / or applications resident on the user device 104. In some embodiments, the user device 104 and / or the content provider device 106 can be designed to use various types of software and / or operating systems. In various illustrative embodiments, the user device 104 and / or the content provider device 106 can be equipped with an input / output device 110 and / or associated with the input / output device 110. For example, the input device can include one or more user input devices (e.g., a keyboard, a mouse, a remote control, a touch screen, etc.). The input / output device 110 can also include one or more display devices (e.g., a television, a monitor, a CRT, a plasma, an LCD, an LED, a touch screen, etc.) or other devices to output information to the user of the user device 104 and / or the user of the content provider device 106.
[0047] The user device 104 and / or the content provider device 106 can be configured to receive data from various sources via the network interface 112 using the network 102. In some embodiments, the network 102 can include a computing network (e.g., a LAN, a WAN, the Internet, etc.), and the network interface 112 of the user device 104 and / or the content provider device 106 can be connected to the computing network via any type of network connection (e.g., a wired network connection (such as Ethernet, a telephone line, a power line, etc.) or a wireless network connection (such as WiFi, WiMAX, 3G, 4G, satellite, etc.)). In some embodiments, the network 102 can include a media distribution network configured to distribute media programs and / or data content, such as a cable (e.g., a coaxial metal cable), a satellite, an optical fiber, etc.
[0048] In some embodiments, the content management system 108 is configured to select third-party content items to be presented on a resource. The content management system 108 includes a processor 120, processing circuitry 118, and a memory 122. Instructions may be stored on the memory 122 that, when executed by the processor 120, cause the processing circuitry 118 to perform the various operations described herein. The operations described herein may be implemented using software, hardware, or a combination thereof. The processor 120 may include a microprocessor, an ASIC, an FPGA, etc., or a combination thereof. In many embodiments, the processor 120 may be a multi-core processor or a processor array. The processor 120 may implement or facilitate a secure environment. For example, the processor 120 may implement software guard extensions (SGX) to define a private region (e.g., an enclave) in the memory 122. The memory 122 may include, but is not limited to, an electronic, optical, magnetic, or any other storage device capable of providing program instructions to the processor 204. The memory 122 may include a floppy disk, a CD-ROM, a DVD, a magnetic disk, a memory chip, a ROM, a RAM, an EEPROM, an EPROM, a flash memory, an optical medium, or any other suitable memory from which the processor 204 may read instructions. The instructions may include code from any suitable computer programming language, such as, but not limited to, C, C++, C#, Java, JavaScript, Perl, HTML, XML, Python, and Visual Basic.
[0049] Memory 122 includes a content database 126 and content analysis circuitry 124. In some embodiments, the content analysis circuitry 124 is configured to conduct an auction or a bidding process. The content analysis circuitry 124 may be configured to select one or more content items 128 of one or more winners of the auction or bidding process for display on a resource. In some embodiments, the content analysis circuitry 124 is further configured to use a quality score (i.e., a measure of the likelihood that a user of the user equipment 104 will interact with the content item 128 or take a conversion action related to the content item 128) or other metrics during the selection process of the content item 128. In some embodiments, a content provider may create a content campaign or may otherwise provide various settings or guidelines to the content management system 108. Such settings or guidelines may govern how the content provider participates in the auction or bidding process (e.g., how much to bid in a given auction, the content provider's total budget (weekly, daily, or otherwise), etc.). Such settings or guidelines may be set based on various metrics, such as cost per impression or cost per thousand impression (CPM), cost per click (CPC), or cost per acquisition (CPA) or cost per conversion. Such settings or guidelines may also be set based on the type of platform on which the content item 128 should be provided (e.g., mobile, desktop, etc.), the type of resource on which the content item 128 should be provided (e.g., search results page), the geographical location of the user equipment on which the resource is displayed, etc. In some embodiments, the settings or guidelines provided to the content management system 108 are stored in the content database 126.
[0050] The query processing system 150 may facilitate an assessment of the accuracy of the manual interpretation of a search query. In various embodiments, the query processing system 150 receives user interaction data from the content provider device 106 and / or from the user equipment 104 of an analysis session. The query processing system 150 may also receive client data and content data of various client providers from the content management system 108. In some embodiments, the query processing system 150 performs various functions on the session data and the client data to generate an estimate of whether the provided manual interpretation of the search query is correct. In various embodiments, the query processing system 150 is a secure environment such that it does not allow access to non-anonymous data. The query processing system 150 may be a server, a distributed processing cluster, a cloud processing system, or any other computing device. The query processing system 150 may include or execute at least one computer program or at least one script. In various embodiments, the query processing system includes a combination of software and hardware, such as one or more processors configured to execute one or more scripts. Referring below to Figure 2Describe the query processing system 150 in more detail.
[0051] Figure 2 FIG. shows a block diagram of a query processing system 150 configured to evaluate whether a human interpretation of a search query is correct according to some embodiments. The query processing system 150 is shown to include a processing circuit 202 having a processor 204 and a memory 206. Instructions may be stored on the memory 206 which, when executed by the processor 204, cause the processing circuit 202 to perform the various operations described herein. The operations described herein may be implemented using software, hardware, or a combination thereof. The processor 204 may include a microprocessor, an ASIC, an FPGA, etc., or a combination thereof. In many embodiments, the processor 204 may be a multi-core processor or a processor array. The processor 204 may implement or facilitate a secure environment. For example, the processor 204 may implement Software Guard Extensions (SGX) to define a private area (e.g., enclave) in the memory 206. The memory 206 may include, but is not limited to, electronic, optical, magnetic, or any other storage device capable of providing program instructions to the processor 204. The memory 206 may include a floppy disk, a CD-ROM, a DVD, a magnetic disk, a memory chip, a ROM, a RAM, an EEPROM, an EPROM, a flash memory, an optical medium, or any other suitable memory from which the processor 204 can read instructions. The instructions may include code from any suitable computer programming language, such as, but not limited to, C, C++, C#, Java, JavaScript, Perl, HTML, XML, Python, and Visual Basic. The memory 206 may include a model training circuit 208, a clustering generation circuit 210, a similarity scoring circuit 212, an interpretation evaluation circuit 214, and a database 220.
[0052] The model training circuit 208 may train a first model for generating an initial evaluation of the accuracy of a human interpretation. In various embodiments, the model training circuit 208 receives a training data set that includes a set of search queries, human interpretations provided for the search queries, and human-labeled values indicating whether the interpretations are accurate. The model training circuit 208 may use the training data set to train the first model. Additionally, the model training circuit 208 may generate and train a final second model by combining the output of the initial evaluation from the first model with one or more other inputs related to the search query. For example, the model training circuit 208 may receive clustering data for the search query from the clustering generation circuit 210 for training the second model. In various embodiments, the model training circuit 208 may utilize the first model as a service for generating token embeddings of the search query. The model training circuit 208 may then average the token embeddings of the search query to generate a fixed vector sentence representation for each search query.
[0053] According to various embodiments, the clustering generation circuit 210 may determine clustering features related to a search query. In addition to determining whether a search query should be grouped with existing clusters, the clustering generation circuit 210 may also create a new cluster for the search query. In some embodiments, the clustering generation circuit 210 is configured to apply clustering techniques (such as K-means clustering, density-based clustering, etc.) to assign search queries to clusters. The clustering generation circuit 210 may control the size of one or more clusters based on whether a "similar search query" cluster or a "broad search query type" cluster is desired. In some embodiments, before assigning a search query that has been stored in the database 220 for a long time, the clustering generation circuit 210 first assigns a more recent search query to a cluster. In this way, the query processing system 150 may prioritize clustering and analyzing more recent search queries rather than search queries issued in the more distant past. In some embodiments, the clustering generation circuit 210 assigns search queries to clusters as service signals added while connected to the network 102. In response to the disconnection of the connection to the network 102 (i.e., offline), the clustering generation circuit 210 may create a dashboard of clusters using the clustering output. Additionally, the clustering generation circuit 210 may be configured to update existing clusters in response to receiving a new search query.
[0054] The similarity scoring circuit 212 may receive a vector sentence representation of a search query and determine a similarity score between two search queries. For example, the similarity scoring circuit 212 may utilize a similarity method (such as the cosine similarity algorithm) on the vector sentence representation to calculate a value representing the degree of similarity between one search query and another search query. Additionally, the similarity scoring circuit 212 may receive a vector sentence representation from the model training circuit 208 of the search query and an associated given interpretation to determine the degree of similarity between the search query and the given interpretation of the search query. The similarity scoring circuit 212 may calculate the degree of similarity between the human interpretation and the corresponding search query on a predetermined scale (e.g., 0 to 1, 0% to 100% similarity, etc.). Then, the interpretation evaluation circuit 214 may use this calculation when evaluating whether the interpretation is correct. In various embodiments, the similarity scoring circuit 212 also parses the search queries to compare similarities. The similarity scoring circuit 212 may then use a distance algorithm (such as the cosine distance algorithm or the Euclidean distance algorithm) to evaluate the similarity between search queries and / or between a search query and a given answer to the search query.
[0055] According to some embodiments, the explanation evaluation circuit 214 includes an initial evaluation circuit 216 and a final evaluation circuit 218. The initial evaluation circuit 216 may utilize a first model trained by the model training circuit 208 to generate a first estimate of whether there is an obvious error in the explanation of a search query. In some embodiments, the initial evaluation circuit 216 may determine the first evaluation based on calculations from the similarity scoring circuit 212 and the clustering generation circuit 210. For example, if the user inputs a search query that the client has historically input over a long period of time, the frequency at which the user continues to input the same question may be an indication of the correct explanation of the search query. The final evaluation circuit 218 may use an overall second model generated and trained by the model training circuit 208 to determine an estimate of whether the human-generated explanation of the search query is correct. In some embodiments, the final evaluation circuit 218 gives an estimate of the overall evaluation by assessing factors such as whether the user has selected a report in the insights card for the answer to the search query and whether additional refined search queries have been made after the original first search query was issued. In various embodiments, the final evaluation circuit 218 outputs the evaluation to one or more user devices 104. The final evaluation circuit 218 may also determine the confidence level of the overall evaluation. For example, the final evaluation circuit 218 may output a 1 or 0 for whether the explanation is accurate, along with a percentage for the determined confidence. In various embodiments, the final evaluation circuit 218 weights factors such as whether a subsequent search refinement query has been issued and whether the user has selected an option to view the report differently depending on the configuration settings of the query processing system 150.
[0056] The database 220 may store and update search queries that can be used by the systems disclosed herein. Additionally, the database 220 may store the generated query clusters. In various embodiments, the search queries do not include dates. For example, if the user inputs a search query such as "How many sessions have occurred since November 2019", the query processing system 150 may strip and disregard the date in the search query. In other embodiments, search queries that include dates are stored in the database 220 and are used to evaluate the accuracy of the human-generated explanation of the search query. The database 220 may include one or more storage media. The storage media may include, but are not limited to, magnetic storage devices, optical storage devices, flash memory, and / or RAM. The query processing system 150 may implement or facilitate various APIs to perform database functions (i.e., manage the data stored in the database 220). The APIs may be, but are not limited to, SQL, ODBC, JDBC, and / or any other data storage device and manipulation APIs.
[0057] In some embodiments, the query processing system 150 may include one or more computing platforms configured to communicate with a remote platform according to a client / server architecture, a peer-to-peer architecture, and / or other architectures. The remote platform may be the same as or similar to the user device 104 described in conjunction with Figure 1 The remote platform may be configured to communicate with other remote platforms via the computing platform and / or according to a client / server architecture, a peer-to-peer architecture, and / or other architectures.
[0058] It should be understood that although the circuits 208, 210, 212, 214, 216, and 218 are shown as being implemented within a single processing unit in Figure 2 In embodiments where the processor 204 includes multiple processing units, one or more of the circuits 208, 210, 212, 214, 216, and 218 may be implemented remotely from the other circuits. The description of the functions provided by the different circuits 208, 210, 212, 214, 216, and 218 is for illustrative purposes and is not intended to be limiting, as any one of the circuits 208, 210, 212, 214, 216, and 218 may provide more or less functionality than described.
[0059] Referring to Figure 3 , a flowchart of a process 300 for evaluating the accuracy of a human interpretation of a search query is shown. The process 300 may be performed by, for example, the query processing system 150 described in conjunction with Figures 1 - 2 In some embodiments, the process 300 is performed in response to the query processing system 150 receiving a new search query.
[0060] Figure 3Process 300 in accordance with one or more embodiments is shown. Process 300 may be performed by one or more components of query processing system 150 in computing environment 100. At 302, process 300 includes receiving a training data set by one or more processing circuits. In some embodiments, model training circuit 208 is configured to perform the operations at 302. In other embodiments, the operations described at 302 may be performed by one or more hardware processors configured by machine-readable instructions, including circuits identical or similar to model training circuit 208. The training data set may include a collection of past search queries, human interpretations of the past search queries, and human evaluation labels as to whether the human interpretations are correct. Thus, the training data set may include three different types of associated data points, namely, past search queries, human interpretations, and human-evaluated labels. In some embodiments, the past search query data may be associated with one or more human interpretations and corresponding human-evaluated labels. For example, the past search query "slowest page" may be associated with the human interpretation value "page with the fewest page views" and the human evaluation label "incorrect" interpretation. Additionally, the past search query "slowest page" may also be associated with the human interpretation value "page with the most average page load time" and the human evaluation label "correct" interpretation in a different collection within the training data set.
[0061] In some examples, each collection in the training data set further includes a score for the results of the search query, a status of the results of the search query, whether the user selected to view a report in the results section of the search query (e.g., an insight card), and / or whether the user selected a different recommended report. Additionally, the collection may include subsequent search queries and human interpretations within a smaller predetermined time period of the labeled search query. For example, the collection may further include subsequent search queries entered by the user after the labeled search query, the amount of time from the first search query to when the subsequent search query was entered, the human interpretation of the subsequent search query, whether the subsequent search query is a refinement of the original first search query, and a similarity score to the original first search query.
[0062] At 304, process 300 includes training a first model by one or more processing circuits using a training data set. The first model can be trained to evaluate whether an artificial interpretation of a search query is correct. In some embodiments, model training circuit 208 is configured to perform the operations described at 304. In other embodiments, the operations performed at 304 can be performed by one or more hardware processors configured by machine-readable instructions, including circuits identical or similar to model training circuit 208. In some embodiments, the first model is a Bidirectional Encoder Representations from Transformers (BERT) model. The first model can initially be pre-trained on a large language data corpus such as Wikipedia. In some embodiments, the operations performed at 304 are used to further train (i.e., fine-tune) the first model to train the first model on information in a training data set from a received specific classifier. The first model can be trained to compute a direct assessment of whether an artificial interpretation of a search query is correct. Additionally, the first model can be used as an embedding layer to compute similarities between search queries within a cluster or different clusters and between a search query and subsequent search queries.
[0063] At 306, process 300 includes receiving, by one or more processing circuits, a search query and an artificial interpretation of the search query. The search query includes one or more words entered via a user interface to request desired information. The artificial interpretation can include one or more words that define an interpretation of the desired information. For example, the search query might be "How many users checked out?" and the artificial interpretation associated with that search query can be "Product checkout quantity". In some embodiments, interpretation evaluation circuit 214 is configured to perform the operations at 306. Interpretation evaluation circuit 214 can receive the search query and the artificial interpretation of the search query from content management system 108 via network 102.
[0064] At 308, process 300 includes an initial assessment by one or more processing circuits of whether a human interpretation of a search query is correct using a first model. According to some embodiments, the operations described at 308 may be performed by an initial assessment circuit 216. The initial assessment circuit 216 may use the first model to evaluate whether a human interpretation of a search query is correct by comparing the search query to a given response. For example, the search query and the given interpretation may be input into the first model, and the first model may output the probability that the interpretation is an incorrect response. The initial assessment circuit 216 may analyze the similarity score of the input search query and the human interpretation calculated by the similarity score circuit 212 to make a first estimate. The initial assessment circuit 216 may also receive from the similarity scoring circuit 212 a comparison of the search query to any potential refinements of the search query to determine whether a subsequent search query is a refinement of the first search query. For example, the similarity scoring circuit 212 may output the similarity between search queries by comparing the average of the token embeddings (e.g., vector representations) of the first query and the subsequent search query from the model training circuit 208. Taking the average of the token embeddings of the search query and the subsequent search query may include average pooling of the penultimate hidden layer of each token in the sentence of the query to achieve a fixed representation of the sentence in the query.
[0065] At 310, process 300 includes generating a second model by one or more processing circuits using an initial evaluation from a first model, time features associated with a search query, and clustering features associated with the search query. In some embodiments, model training circuit 208 is configured to perform the operations performed at 310. In some embodiments, the second model generated and trained by model training circuit 208 is a logistic regression model. Model training circuit 208 may receive from initial evaluation circuit 216 an initial evaluation of whether a human interpretation of the search query is correct. In various embodiments, model training circuit 208 receives from the first model an initial evaluation of whether the human interpretation is correct, along with session data, client data, and user data. Session data may include the duration of the session, the number of search queries input during the session, the timing of each search query during the session, user interactions with the results of the search queries during the session (e.g., selecting to view a report), etc. Client data may include information related to the content provider, such as the most frequently input search queries for the client, data about the content provided by the client, past analysis results for the client, etc. In some embodiments, user data includes information about a particular user of the content provider device (e.g., content provider device 106), such as past search queries input by the user, the user's average session time, etc. According to some embodiments, data provided as input to the second model to determine a final evaluation may be retrieved from database 220. In other embodiments, model training circuit 208 may receive inputs for the final second model from content management system 108 via network 102.
[0066] Clustering features related to a search query can include the cluster closest to the search query, whether a subsequent second search query is in the cluster containing the search query, user data associated with search queries in the cluster containing the search query, or session data associated with search queries in the cluster containing the search query. For example, the model training circuit 208 can be configured to train a second model to analyze the average number of different clusters (e.g., the diversity of clusters) that a user inputs search queries against during a session, changes in cluster membership in terms of verticals, countries, etc. (e.g., the segmentation of clusters), and the largest cluster in terms of size among search queries with predicted significant errors (e.g., search queries that are most frequently answered incorrectly). Additionally, the cluster with the lowest answer coverage among all search queries (e.g., search query growth), whether subsequent search queries issued by the user are in the same cluster as the previous search query (e.g., search refinement), the cluster closest to the search query (e.g., alternative interpretations), and which search queries should be rated (e.g., sampling efficiency) can also be analyzed by the model training circuit 208 when training the second model. The second model generated and trained by the model training circuit 208 can utilize these clustering features related to the search query to determine whether the manual interpretation of the search query is accurate.
[0067] In some embodiments, time features related to a search query include the frequency of issuing the search query within a past time period, the frequency of search queries in the cluster where the search query is issued, or the frequency of suggesting the search query as an input via the user interface. For example, time features related to a search query can include the time between when a user makes a search query and other search queries issued within a predetermined time period before or after the search query. Time features related to a search query can also include whether the search query was made by the user in a previous user session. For example, the model training circuit 208 can analyze whether a search query is a historical search query of the user stored in the database 220. Additionally, the second model can be trained by the model training circuit 208 to consider the average frequency of search queries issued by the user in the past and the average frequency of search queries in the same cluster as the search query issued by the user in the past. In some embodiments, the model training circuit 208 uses additional inputs to generate and train the second model to make a final assessment of the accuracy of the manual interpretation of the search query.
[0068] At 312, process 300 includes a final evaluation by one or more processing circuits of whether a human interpretation of a search query is correct using a second model. In some embodiments, the operations performed at 312 may be performed by final evaluation circuit 218. Final evaluation circuit 218 may use the second model to generate a final evaluation of whether the human interpretation is correct based on a similarity score received from similarity scoring circuit 212 for a subsequent search query issued by a user. Additionally, final evaluation circuit 218 may treat reported user interactions in the result area for the search query as an indication of whether the human interpretation is correct, which is described in more detail below. In some embodiments, the value of the final evaluation is 0 for a correct interpretation or 1 for an incorrect interpretation. The final evaluation determined by the second model via final evaluation circuit 318 may be associated with a confidence score, such as a percentage value for the confidence that the second model did not make an error in its final evaluation. In other embodiments, final evaluation circuit 218 is configured to use the second model to generate a final evaluation that is the probability that a human interpretation of a search query is correct. For example, final evaluation circuit 218 may determine an output of 91% from the second model (i.e., the human interpretation of the search query is an accurate answer).
[0069] Now referring Figure 4 , a flowchart of a process 400 for analyzing a search query to evaluate whether a human interpretation of the search query is an accurate interpretation is shown. Process 400 may be performed by, for example, query processing system 150 described with reference to Figure 1 and Figure 2 . Figure 4 FIG. shows process 400 in accordance with one or more embodiments. In some embodiments, process 400 is performed by query processing system 150 during the operations performed at 312 during process 300( Figure 3 ).
[0070] According to some embodiments, at 402, process 400 includes determining by one or more processing circuits whether a first search query is associated with a second search query received within a predetermined time interval. Clustering generation circuit 210 may be configured to perform the operations performed at 402. Clustering generation circuit 210 may determine whether the user has entered an additional search query from content provider device 106 after the original search query. For example, the search query with a human interpretation whose accuracy is being evaluated may be a first question from the user, and the second search query may be a subsequent question from the user. In some embodiments, clustering generation circuit 210 may receive session data related to the subsequent search query from database 220.
[0071] According to some embodiments, at 404, process 404 includes receiving, by one or more processing circuits, token embeddings from a first model for a first search query and a second search query. According to some embodiments, the operations described at 404 may be performed by model training circuit 208. Model training circuit 208 may be configured to use the first model to generate token embeddings for each search query analyzed by query processing system 150. A token is a word in a search query. An embedding is a mathematical representation of a word, such as a vector representation. In this way, similar words have a smaller distance between their respective embeddings. For example, the distance from the token embedding of the word "important" to the token embedding of the word "significant" is smaller than the distance to the token embedding of the word "excellent". Model training circuit 208 may then send the token embeddings for the search queries to similarity scoring circuit 212 to evaluate how similar the search query is to another search query, or how similar the search query is to an interpretation of the search query.
[0072] According to some embodiments, at 406, process 400 includes determining, by one or more processing circuits, vector sentence representations for the first search query and the second search query. The operations performed at 406 may be carried out by model training circuit 208. Model training circuit 208 may take an average of the token embeddings from the first model for each of the first search query and the second search query to generate a vector sentence representation for each of the first search query and the second search query. For example, the first model may use average pooling of the token embeddings (i.e., vector representations) for each search query to determine a fixed representation of the sentence in the search query. Then, similarity scoring circuit 212 may utilize the vector sentence representations for the search queries to generate a similarity score between the search queries.
[0073] According to some embodiments, at 408, process 400 includes generating, by one or more processing circuits, a similarity score for the first search query and the second search query. Similarity scoring circuit 212 may be configured to perform the operations performed at 408. The similarity score may be generated by the first model using a similarity-based metric (e.g., cosine similarity calculation, Jaccard similarity calculation, etc.) on the vector sentence representations. In some embodiments, the similarity score ranges from a value of 0 to 1, where 0 occurs when the vector sentence representations are the same. In other embodiments, the similarity score may be generated by similarity scoring circuit 212 on a different scale or as a percentage value. In some embodiments, similarity scoring circuit 212 may also be configured to parse the first search query and the second search query using the first model to determine the similarity between the search queries using a distance algorithm (such as a cosine distance algorithm or an Euclidean distance algorithm).
[0074] Now turning to Figure 5 , a flowchart of a process 500 for evaluating the accuracy of a human interpretation of a search query according to an exemplary embodiment is shown. The process 500 may be performed by, for example, the query processing system 150 described with reference to Figures 1 - 2 . The process 500 may be performed in response to an operation at 408 (refer to Figure 4 ) that is being executed. At 502, the model training circuit 208 may determine whether a second search query is a search refinement of a first search query. The model training circuit 208 may infer whether a subsequent search query is a search refinement of the first search query based on the value of the similarity score between the search queries from the similarity scoring circuit 212. For example, if the similarity score calculated at 408 in the process 400 is below a predetermined threshold, the model training circuit 208 may determine that the second search query is not a search refinement of the original search query. If the model training circuit 208 determines at 502 that the second search query is not a search refinement, the process 500 proceeds to 504. According to some embodiments, at 504, the final evaluation circuit 218 determines a weighted indication of the correct human interpretation of the search query. For example, if two vector sentence representations of a search query have a high similarity score and the indication vectors are very different, the final evaluation circuit 218 may evaluate that the second search query is an unrelated follow-up question and thus indicate that the interpretation of the first search query is correct.
[0075] However, if at 502 the model training circuit 208 determines that the second search query is a search refinement of the first search query, the process 500 proceeds to 506. For example, if the similarity score between the first search query and the subsequent search query is higher than a predetermined threshold (e.g., 0.8, 75%, etc.), the second model determines via the model training circuit 208 that the subsequent search query is a search refinement of the first search query, and the process 500 proceeds to 506. According to some embodiments, at 506, the final evaluation circuit 218 determines a weighted indication of an incorrect human interpretation of the search query. For example, if the first search query is "slowest page" and its interpretation is "page with the fewest views", and the subsequent search query is "page that loads the slowest", the final evaluation circuit 218 determines that the interpretation of "page with the fewest views" may be a weighted indication of an incorrect interpretation. This indication is a weighted feature input to the second model when determining the overall evaluation. In some embodiments, the final evaluation circuit 208 is configured to train the second model to weight the indication determined at 502 by a predetermined amount when evaluating whether the interpretation of the search query is correct and other factors. In this way, the final evaluation of whether the interpretation of the search query is correct is based on the trained weighting of this indication and other feature inputs (e.g., historical search queries and human interpretation data).
[0076] According to some embodiments, at 508, process 500 may include determining whether an input is received for viewing a report on a search query. The final evaluation circuit 218 may be configured to receive session data from the database 220, the session data including information about user interactions during the session in which the search query was issued. For example, the final evaluation circuit 218 may receive data indicating that the user selected an option to view a report in the result area of the search query via a user interface of the content provider device 106 or the user device 104 (e.g., user interface 900( Figure 9 ))). In some embodiments, the final evaluation circuit 218 may alternatively receive session data from the clustering generation circuit 210 or the model training circuit 208. In some embodiments, the final evaluation circuit 218 may also be configured to determine at 412 whether the user has selected an option to view other suggested reports in the result area of the search query.
[0077] If the final evaluation circuit 218 determines at 508 that a user input is received for viewing a report displayed in the result area for the search query, process 500 proceeds to 510. According to some embodiments, at 510, the final evaluation circuit 218 determines a weighted indication of the correct human interpretation of the search query. When evaluating whether the interpretation is correct, the final evaluation circuit 218 may place a greater weight on the indication from the user input to view the report in the result area than on other factors. In other embodiments, the final evaluation circuit 218 may place a greater weight on whether the user inputs a subsequent search query that is a refinement of the first search query.
[0078] On the other hand, if at 508 the final evaluation circuit 218 determines that no user input is received for viewing a report on the search query, process 500 proceeds to 512. At 512, the final evaluation circuit 218 may determine a weighted indication of an incorrect human interpretation of the search query. In some embodiments, the indications determined at 510 and 512 may not have an equal impact on the final evaluation of whether the interpretation made by the second model is accurate. For example, receiving a user input to view a report may have a greater impact on the final evaluation made by the final evaluation circuit 218 than not receiving a user input (i.e., the second model may be trained to place a greater weight).
[0079] Now referring to Figure 6 , a flowchart of a process 600 for evaluating the accuracy of a human interpretation of a search query using clustering according to some embodiments is shown. The query processing system 150 described with reference to Figure 1 and Figure 2 may execute process 600. In some embodiments, process 600 is performed by the query processing system 150 in the execution of reference Figure 3It is performed after the operations at 302 and 304. At 602, process 600 includes generating clusters based at least in part on the similarity between search queries in a cluster of search queries. In some embodiments, the cluster generation circuit 210 may be configured to receive a vector sentence representation of search queries generated by a first model from the model training circuit 208. The cluster generation circuit 210 may then assign the search queries to clusters using a given clustering technique. In some embodiments, the clustering technique uses a larger k value to cluster semantically similar search queries.
[0080] According to some embodiments, at 604, the cluster generation circuit 210 may be configured to receive a new search query. For example, the cluster generation circuit 210 may receive a new search query from the database 220, from another component of the processing circuit 202, or from the content management system 108. In some embodiments, the new search query is a past search query that is being classified into a cluster. At 606, process 600 includes updating the generated clusters based on the similarity between the search queries in the cluster and the new search query. The cluster generation circuit 210 may be configured to classify the new search query into the cluster based on a similarity score with other search queries in the cluster (e.g., as described with reference to Figure 4 that described at 408). In some embodiments, in addition to assigning incoming search queries and corresponding explanations to existing clusters, the cluster generation circuit 210 is configured to search for and discover new clusters. The cluster generation circuit 210 may use the cluster output to generate a cluster dashboard to allow online and offline access to the clusters. In some embodiments, when disconnected from the network 102, the cluster generation circuit 210 creates a dashboard of the clusters. However, when connected to the network 102, the cluster generation circuit 210 may continuously assign search queries to clusters as an additional service signal.
[0081] In some embodiments, the created dashboard depicts various classifications of errors automatically generated by the overall model via the cluster generation circuit 210. Advantageously, the methods disclosed herein may allow a user to view which search queries have the highest percentage of human interpretation errors and related search queries. For example, the clusters determined by the cluster generation circuit 210 are used for reporting purposes, such as reporting that a search has a certain percentage of human interpretation errors for search queries in cluster A. Cluster A may be characterized by search queries that are similar to the top ten most popular search queries in the cluster. In this way, information about which human interpretations are generating the most errors can be determined. Typically, humans are used to manually evaluate how to classify different search queries to generate statistics about the search queries and to report details about human interpretation errors. Advantageously, the systems and methods disclosed herein facilitate the automatic classification of errors in search queries and human interpretations.
[0082] Referring now to Figure 7 , there is shown a diagram 700 of a high-level modeling process of an overall model used by the query processing system described in Figure 1 and Figure 2 . Diagram 700 shows the input and output of an overall model for evaluating the accuracy of an artificial interpretation of a natural language search query utilized by query processing system 150. The input includes search query 702, interpretation 704, session data 710, user data 712, and client data 714. Search query 702 may be input via a single search via a user interface. A single search is a search bar where a user inputs a query in an analytics application. Interpretation 704 is an artificial interpretation of search query 702, meaning that interpretation 704 is a human language interpretation of what the user wants to be displayed as a result in an insight card (e.g., Figure 9 's insight card 908) on the user interface. The methods disclosed herein focus on evaluating whether there is an obvious error (i.e., interpretation 704 is an incorrect interpretation of search query 702) for search query 702 and interpretation 704. Search query 702 and interpretation 704 are used as input to a first model 706.
[0083] In some embodiments, the first model 706 is a BERT model. The first model 706 may support fine-tuning after the first model 706 is pre-trained on a large language data corpus and then specialized on a smaller manually rated language data corpus for a specific classification problem. This is beneficial because the first model 706 requires less manually rated data for manual evaluation subsequently, and collecting such data is expensive. Additionally, fine-tuning the first model 706 can beneficially improve efficiency and the system's bandwidth because a smaller set of human-labeled training sets can be used to specialize the first model 706. Session data 710, user data 712, and client data 714, as well as the output of the initial evaluation of whether the interpretation is accurate from the first model, are used as input to a final model 708. Session data 710 may include information about other search queries input during a session, the frequency of the input search query, etc. User data 712 may include statistical data of a particular user, such as the average amount of search queries issued by the user during a session. In some embodiments, client data 714 may include information about the overall client (e.g., a content provider company), such as the type of content provided by the client, the most common type of search query input for the client, etc. The final model 708 may be a combined model that uses the first model 706 in addition to various clustering features and temporal features related to search query 702 to identify an evaluation 716 of the correct interpretation. The evaluation 716 of the correct interpretation is an estimate of whether interpretation 704 of search query 702 is accurate or whether there is an obvious error for interpretation 704.
[0084] Now turning to Figure 8 , a model architecture of a first model 706 and use cases of the first model 706 according to some embodiments are shown. Figure 8 Diagram 800 showing the architecture of the first model 706 and diagram 850 of specific use cases of the first model 706 used in the systems and methods disclosed herein. Diagram 800 shows the bidirectional transformer architecture of the first model 706. E k represents a token (roughly speaking, a word), and T k represents a digital vector, which is a representation of each input token in the (bidirectional) context. T rm represents a transformer unit. The first model 706, which is a BERT model, can jointly learn from information about the left and right contexts in all layers, encode words into vectors, and output digital vectors corresponding to the words. The first model 706 can utilize a mechanism called self-attention, which is contrary to the sequence mechanism in, for example, a recurrent neural network (RNN). For a transformer unit, each word can be assigned 3 vectors, a key, a value, and a query. The elements of each of these vectors are learned by the first model 706. Then, each word can "query" each other word by calculating the dot product between its query and the keys of other words. This may then result in a weight for each word for each query. For each query, the transformer unit returns a weighted average of the value vectors. Thus, the output of each transformer unit is a vector for each word or query, which is then passed to a feed-forward neural network. Then, this process is repeated several times depending on the size of the first model 706.
[0085] Diagram 850 of the use cases of the first model 706 depicts the fine-tuning (i.e., transfer learning) ability of the first model 706. Special tokens are used in the use case of sentence classification, which is useful if there is a downstream task for the first model 706 to learn. Diagram 850 shows the combination of the first model 706 (e.g., BERT) with an additional output layer. Advantageously, the model then starts learning with a minimal number of parameters. In diagram 850, E represents the input embedding, T i represents the context representation of token i, [CLS] represents a special symbol for classification output, and [SEP] is a special symbol for separating non-consecutive token sequences. The fine-tuning of the first model 706 pools the data in a single output token. Once the output token is created, the final model can use the vector representation of the classification output as input and be trained with the corresponding labels.
[0086] Now referring to Figure 9, which shows a user interface for entering a search query and viewing the human interpretation and result data for the search query according to some embodiments. The user interface 900 is shown to include a single search 902, a human interpretation 904, optional report options 906, and an insight card 908. The single search 902 can be a search bar for the user to enter a search query to view analysis data on the content provided by the content management system 108. In some embodiments, the single search 902 automatically populates suggestions for the search query based on common search queries entered by the user in past sessions. In this exemplary embodiment, the entered search query is "countries with the most users", and the corresponding human interpretation 904 is "countries with the most users" as shown in the insight card 908. The optional report options 906 can allow the user to view the generated report by selecting a link displayed in the insight card 908. The insight card 908 is a result area displayed to the user in the right panel, where insight results are shown for the entered search query. In some embodiments, the insight card 908 displays options for viewing other suggested reports. The insight card 908 can also allow the user to select an option for asking follow-up questions (i.e., entering a follow-up search query).
[0087] Figure 10 Depiction of a computing system 1000 is shown, which can be used to implement, for example, the illustrative user device 104, the illustrative content management system 108, the illustrative content provider device 106, the illustrative query processing system 150, and / or various other illustrative systems described in the present disclosure. The computing system 1000 includes a bus 1008 or other communication components for conveying information, and a processor 1012 coupled to the bus 1008 for processing information. The computing system 1000 also includes a main memory 1002 coupled to the bus 1008 for storing information and instructions to be executed by the processor 1012, such as a random access memory (RAM) or other dynamic storage device. The main memory 1002 can also be used to store location information, temporary variables, or other intermediate information during the execution of instructions by the processor 1012. The computing system 1000 can also include a read only memory (ROM) 1004 or other static storage device coupled to the bus 1008 for storing static information and instructions for the processor 1012. A storage device 1006, such as a solid state device, a magnetic disk, or an optical disk, is coupled to the bus 1008 for permanently storing information and instructions.
[0088] The computing system 1000 can be coupled via a bus 1008 to a display 1014, such as a liquid crystal display or an active matrix display, for displaying information to a user. An input device 1016, such as a keyboard including alphanumeric keys and other keys, can be coupled to the bus 1008 for transmitting information and command selections to the processor 1012. In another embodiment, the input device 1016 has a touchscreen display 1014. The input device 1016 can include a cursor control, such as a mouse, a trackball, or cursor direction keys, for transmitting direction information and command selections to the processor 1012 and for controlling cursor movement on the display 1014.
[0089] In some embodiments, the computing system 1000 can include a communication adapter 1010, such as a network adapter. The communication adapter 1010 can be coupled to the bus 1008 and can be configured to enable communication with a computing or communication network 1018 and / or other computing systems. In various illustrative embodiments, any type of networking configuration can be implemented using the communication adapter 1010, such as wired (e.g., ), wireless (e.g., via etc.), pre-configured, ad-hoc, LAN, WAN, etc.
[0090] In accordance with various embodiments, the processes implementing the illustrative embodiments described herein can be implemented by the computing system 1000 in response to an instruction arrangement contained in the main memory 1002 being executed by the processor 1012. Such instructions can be read into the main memory 1002 from another computer-readable medium, such as the storage device 1006. Execution of the instruction arrangement contained in the main memory 1002 causes the computing system 1000 to perform the illustrative processes described herein. One or more processors in a multiprocessing arrangement can also be used to execute the instructions contained in the main memory 1002. In alternative embodiments, hardwired circuitry can be used in place of or in combination with software instructions to implement the illustrative embodiments. Accordingly, embodiments are not limited to any specific combination of hardware circuitry and software.
[0091] The systems and methods described in this disclosure can be implemented for any type of third-party content item (i.e., for any type of content item to be displayed on a resource). In one embodiment, the content item can include an advertisement. In one embodiment, the content item can include any text, image, video, story (e.g., a news story), social media content, link, or any other type of content provided by a third party for display on a first-party content provider's resource. The type of content item for which the content visibility methods herein are used is not restrictive.
[0092] Although in Figure 10An example processing system is described, but implementations of the subject matter and functional operations described in this specification can be implemented using other types of digital electronic circuitry, or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents), or in combinations of one or more of them.
[0093] Implementations of the subject matter and operations described in this specification can be implemented using digital electronic circuitry, or in computer software embodied on a tangible medium, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents), or in combinations of one or more of them. Implementations of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more subsystems of computer program instructions, encoded on one or more computer storage media for execution by, or to control the operation of, a data processing apparatus. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device for execution by the data processing apparatus. A computer storage medium can be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them, or be included in one of them. Moreover, although a computer storage medium is not a propagated signal, a computer storage medium can be the source or destination of computer program instructions encoded in an artificially generated propagated signal. A computer storage medium can also be one or more separate components or media (e.g., multiple CDs, disks, or other storage devices), or be included in one of them. Accordingly, a computer storage medium is both tangible and non-transitory.
[0094] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0095] The term "data processing apparatus" or "computing device" encompasses all kinds of apparatus, devices and machines for processing data, including, for example, programmable processors, computers, system-on-a-chip, or multiple ones or combinations of the foregoing. The apparatus may include dedicated logic circuitry, such as a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program being discussed, for example, code that constitutes processor firmware, protocol stack, database management system, operating system, cross-platform runtime environment, virtual machine, or a combination of one or more of them. The apparatus and the execution environment may implement various different computing model infrastructures, such as network services, distributed computing, and grid computing infrastructures.
[0096] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may or may not correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (for example, one or more scripts stored in a markup language document), stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (for example, files that store one or more subsystems, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0097] The processes and logical flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logical flows can also be performed by dedicated logic circuitry, and the apparatus can also be implemented as dedicated logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0098] For example, a processor suitable for executing a computer program includes any one or more processors of general and special purpose microprocessors as well as any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for performing operations in accordance with instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include or be operably coupled to one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, from which to receive data, or to which to transfer data, or both. However, a computer need not have such devices. In addition, a computer may be embedded in another device, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), etc. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, by way of example, including: semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM disks and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0099] To provide for interaction with a user, embodiments of the subject matter described in this specification may be implemented using a computer having a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user may provide input to the computer. Other kinds of devices may also be used to provide for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input received from the user may be in any form, including acoustic, speech, or tactile input. In addition, a computer may interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page to a web browser on a client device of the user in response to a request received from the web browser.
[0100] Embodiments of the subject matter described in this specification can be implemented using a computing system that includes a backend component (e.g., as a data server), or a middleware component (e.g., an application server), or a frontend component (e.g., a client computer having a graphical user interface or a web browser, by which a user can interact with embodiments of the subject matter described in this specification), or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network).
[0101] The computing system can include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. The relationship between the client and the server arises from computer programs that run on respective computers and have a client-server relationship with each other. In some embodiments, the server sends data (e.g., an HTML page) to a client device (e.g., to display data to and receive user input from a user interacting with the client device). Data generated at the client device (e.g., the result of a user interaction) can be received at the server from the client device.
[0102] In some illustrative embodiments, the features disclosed herein may be implemented on a smart TV module (or a connected TV module, a hybrid TV module, etc.), which may include processing circuitry configured to integrate an Internet connection with more traditional TV program sources (e.g., via cable, satellite, over-the-air, or other signal reception). The smart TV module may be physically incorporated into a television set or may include a separate device, such as a set-top box, a Blu-ray or other digital media player, a gaming console, a hotel TV system, and other companion devices. The smart TV module may be configured to allow a viewer to search for and find videos, movies, photos, and other content on the network, on local cable TV channels, on satellite TV channels, or on a local hard drive. A set-top box (STB, set-top box) or a set-top unit (STU, set-top unit) may include an information appliance device that may contain a tuner and be connected to a television set and an external signal source, so as to convert the signal into content and then display the content on a TV screen or other display device. The smart TV module may be configured to provide a home screen or a top-level screen that includes icons for a plurality of different applications (such as a web browser and a plurality of streaming services, connected cable or satellite media sources, other network "channels", etc.). The smart TV module may also be configured to provide an electronic program guide to a user. Companion applications of the smart TV module may operate on a mobile computing device to provide additional information about available programs to the user, thereby allowing the user to control the smart TV module, etc. In alternative embodiments, these features may be implemented on a laptop computer or other personal computer, a smart phone, other mobile phones, a handheld computer, a tablet PC, or other computing devices.
[0103] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or the scope that may be claimed, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described herein in the context of separate embodiments may also be implemented combinatorially or in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments, separately or in any suitable sub-combination. Additionally, although features may be described above as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination may be deleted from that combination, and the claimed combination may be directed to a sub-combination or a variant of a sub-combination. Further, features described with respect to a particular heading may be used and / or combined with respect to the illustrative embodiments described under other headings; the headings (if provided) are included only for purposes of readability and should not be construed as limiting any features provided under those headings.
[0104] Similarly, although the operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to obtain the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system components in the above-described embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products contained on a tangible medium.
[0105] Accordingly, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still obtain the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to obtain the desired result. In some embodiments, multitasking and parallel processing may be advantageous.
Claims
1. A method for evaluating the accuracy of a human interpretation of a search query, the method comprising: receiving, by one or more processing circuits, a training data set including a plurality of past search queries, a human interpretation for each of the plurality of past search queries, and a human evaluation label indicating whether the human interpretation for each of the plurality of past search queries is correct; training, by the one or more processing circuits, a first model using the training data set, wherein the first model is trained to evaluate whether a human interpretation of a search query is correct; receiving, by the one or more processing circuits, a search query and a human interpretation of the search query, the search query including one or more words input via a user interface to request desired information, and the human interpretation including one or more words defining an interpretation of the desired information; determining, by the one or more processing circuits, an initial evaluation of whether the human interpretation of the search query is correct using the first model; generating, by the one or more processing circuits, a second model using the initial evaluation from the first model, time features associated with the search query, and clustering features associated with the search query, wherein the clustering features are the search queries assigned to a predetermined cluster; and determining, by the one or more processing circuits, a final evaluation of whether the human interpretation of the search query is correct using the second model.
2. The method according to claim 1, wherein, the search query is a first search query, and wherein the method further comprises: determining, by the one or more processing circuits, whether the first search query is associated with a second search query received within a predetermined time interval after the first search query; receiving, by the one or more processing circuits, token embeddings from the first model for each of the first search query and the second search query, wherein a token is a word in the search query; and determining, by the one or more processing circuits, a vector sentence representation for each of the first search query and the second search query by averaging the token embeddings from the first model for each of the first search query and the second search query.
3. The method according to claim 2, wherein, the first model is pre-trained on a natural language data set, and wherein training the first model using the training data set tunes the first model to a specific classification problem.
4. The method according to claim 2, further comprising: parsing, by the one or more processing circuits, the first search query and the second search query using a distance algorithm, wherein the distance algorithm is at least one of an Euclidean distance algorithm or a cosine similarity algorithm.
5. The method according to claim 2, further comprising: determining, by the one or more processing circuits, whether the second search query is a search refinement of the first search query, wherein a search refinement is a weighted indication of an incorrect human interpretation of the search query.
6. The method according to claim 1, further comprising: generating, by the one or more processing circuits, the clusters based at least in part on similarities between search queries in clusters of search queries of different sizes; and updating, by the one or more processing circuits, the generated clusters in response to receiving a new search query.
7. The method according to claim 1, further comprising: determining, by the one or more processing circuits, whether an input for viewing a report on the search query is received via the user interface, wherein the input for viewing the report is a weighted indication of a correct human interpretation of the search query.
8. A system configured to evaluate the accuracy of a human interpretation of a search query, the system comprising: one or more hardware processors configured by machine-readable instructions to: receive, by one or more processing circuits, a training data set including a plurality of past search queries, a human interpretation of each of the plurality of past search queries, and a human evaluation label indicating whether the human interpretation of each of the plurality of past search queries is correct; train, by the one or more processing circuits, a first model using the training data set, wherein the first model is trained to evaluate whether a human interpretation of a search query is correct; receive, by the one or more processing circuits, a search query and a human interpretation of the search query, the search query including one or more words input via a user interface to request desired information, and the human interpretation including one or more words defining an interpretation of the desired information; determine, by the one or more processing circuits, an initial evaluation of whether the human interpretation of the search query is correct using the first model; generate, by the one or more processing circuits, a second model using the initial evaluation from the first model, time features associated with the search query, and cluster features associated with the search query, wherein the cluster features are search queries assigned to a predetermined cluster; and determine, by the one or more processing circuits, a final evaluation of whether the human interpretation of the search query is correct using the second model.
9. The system according to claim 8, wherein the search query is a first search query, and wherein the one or more hardware processors are further configured by machine-readable instructions to: determine, by the one or more processing circuits, whether the first search query is associated with a second search query received within a predetermined time interval after the first search query; receive, by the one or more processing circuits, token embeddings from the first model for each of the first search query and the second search query, wherein a token is a word in the search query; and determine, by the one or more processing circuits, a vector sentence representation for each of the first search query and the second search query by averaging the token embeddings from the first model for each of the first search query and the second search query.
10. The system according to claim 9, wherein The first model is pre-trained on a natural language dataset, and wherein, the training dataset is used to train the first model to tune the first model to a specific classification problem.
11. The system according to claim 9, wherein, the one or more hardware processors are further configured by machine-readable instructions to: parse the first search query and the second search query by the one or more processing circuits using a distance algorithm, wherein the distance algorithm is at least one of an Euclidean distance algorithm or a cosine similarity algorithm.
12. The system according to claim 9, wherein, the one or more hardware processors are further configured by machine-readable instructions to: determine by the one or more processing circuits whether the second search query is a search refinement of the first search query, wherein a search refinement is a weighted indication of an incorrect human interpretation of the search query.
13. The system according to claim 8, wherein, the one or more hardware processors are further configured by machine-readable instructions to: generate the cluster by the one or more processing circuits based at least in part on similarities between search queries in a cluster of search queries of different sizes; update the generated cluster by the one or more processing circuits in response to receiving a new search query.
14. The system according to claim 8, wherein, the one or more hardware processors are further configured by machine-readable instructions to: determine by the one or more processing circuits whether an input for viewing a report for the search query is received via the user interface, wherein the input for viewing the report is a weighted indication of a correct human interpretation of the search query.
15. A non-transitory computer-readable storage medium having instructions thereon that are executable by one or more processors to perform operations for evaluating the accuracy of a human interpretation of a natural language search query, the operations comprising: receiving, by one or more processing circuits, a training dataset that includes a plurality of past search queries, a human interpretation of each of the plurality of past search queries, and a human evaluation label as to whether the human interpretation of each of the plurality of past search queries is correct; training, by the one or more processing circuits, a first model using the training dataset, wherein the first model is trained to evaluate whether a human interpretation of a search query is correct; receiving, by the one or more processing circuits, a search query and a human interpretation of the search query, the search query including one or more words entered via a user interface to request desired information, the human interpretation including one or more words defining an interpretation of the desired information; determining, by the one or more processing circuits, an initial evaluation of whether the human interpretation of the search query is correct using the first model. The one or more processing circuits generate a second model using an initial evaluation from the first model, temporal features associated with the search query, and clustering features associated with the search query, where the clustering features are the search queries assigned to a predetermined cluster; and The one or more processing circuits use the second model to determine a final evaluation of whether the human interpretation of the search query is correct.
16. The computer-readable storage medium according to claim 15, wherein, the search query is a first search query; wherein the operation is for the one or more processing circuits to determine whether the first search query is associated with a second search query received within a predetermined time interval after the first search query; wherein the operation is for the one or more processing circuits to receive token embeddings from the first model for each of the first search query and the second search query, where a token is a word in the search query; and wherein the operation is for the one or more processing circuits to determine a vector sentence representation for each of the first search query and the second search query by averaging the token embeddings from the first model for each of the first search query and the second search query.
17. The computer-readable storage medium according to claim 16, wherein, the first model is pre-trained on a natural language data set, and wherein the first model is tuned to a specific classification problem using the training data set.
18. The computer-readable storage medium according to claim 16, wherein, the operation further includes: the one or more processing circuits parse the first search query and the second search query using a distance algorithm, where the distance algorithm is at least one of an Euclidean distance algorithm or a cosine similarity algorithm.
19. The computer-readable storage medium according to claim 16, wherein, the operation further includes: the one or more processing circuits determine whether the second search query is a search refinement of the first search query, where a search refinement is a weighted indication of an incorrect human interpretation of the search query.
20. The computer-readable storage medium according to claim 15, wherein, the operation further includes: the one or more processing circuits generate the cluster at least partially based on similarities between search queries in clusters of search queries of different sizes; and the one or more processing circuits update the generated cluster in response to receiving a new search query.
Citation Information
Patent Citations
Cognitive Interactive Searching Method And System Based On Personalized User Model And Context
CN105760417A
Method and system for classification of user query intent for medical information retrieval system
CN107301195A