Generating search lexical elements from queries using language models
By training the language model to predict query generalizations associated with digital components and identifying search terms from these generalizations, the challenges existing in existing systems when finding relevant keywords are solved, efficient query annotation and keyword recognition are achieved, and the complexity and resource consumption of system updates are reduced.
Patent Information
- Application Number
- CN202480004400.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-01
- Filing Date
- 2024-05-28
- Publication Date
- 2025-05-30
AI Technical Summary
Existing systems have challenges in finding keywords associated with digital components, especially when dealing with a large number of queries, and updating these systems can be time-consuming and difficult.
By obtaining multiple training samples, each sample including one or more queries, the language model is trained to generate a trained language model, used to predict the generalization of each training sample, and to determine the search lexicon from these generalizations, and finally the query is associated with its corresponding search lexicon and stored.
Eliminates the need to periodically readjust and regulate the relevance-based generalization process, reduces or even eliminates the possibility of hallucination, and saves processing resources, because once the model is trained, the corpus that can be used to process queries generates search term elements.
Smart Images

Figure CN120077370A_ABST
Abstract
Description
Background Art
[0001] The specification relates to data processing, and in particular, to using a language model to generate retrieval tokens.
[0002] Digital components that are discrete units of digital content or digital information are selected for inclusion in digital content that serves a requesting user device. To select digital components to be served, the system needs to be able to evaluate which digital components are most suitable for serving for a particular service instance. One way to select digital components is to use keywords associated with the digital components.
[0003] However, finding keywords associated with digital components can be challenging. Keywords are typically based on digital content, such as queries that the service system has received. There are usually billions of queries to choose from. In addition, some systems generalize queries, such that providers of digital components or a recommendation process that can identify keywords associated with digital components can search the queries.
[0004] To be able to search these queries efficiently, the queries are annotated with retrieval tokens. Determining which retrieval tokens to use to annotate a particular query can be challenging. Systems that make such determinations are typically processing systems based on relevance and traffic, which can process many different information sources, such as URL text, web resource text, etc. In addition, updating these systems can be time-consuming and difficult. Summary of the Invention
[0005] In general, one innovative aspect of the subject matter described in this specification can be embodied in a method that includes the following actions: obtaining a plurality of training samples, each training sample including one or more queries, each of the one or more queries being a query selected from a query repository; training a language model to generate a trained language model, wherein the trained language model predicts, for each training sample, one or more generalizations that describe the training sample, each of the one or more generalizations being an n-gram; processing, by the trained language model, queries stored in the query repository to predict, for each query, one or more generalizations of the query; for each query, determining, from the one or more generalizations predicted for the query, one or more retrieval tokens for the query; and for each query, storing, in a data repository, an association between the query and the one or more retrieval tokens determined for the query. Other embodiments of this aspect include corresponding systems, devices, and computer programs configured to perform the actions of the method, encoded on a computer storage device.
[0006] These and other embodiments may each optionally include one or more of the following aspects. In one aspect, training a language model to generate a trained language model includes: for each training sample, training the model to predict a seed keyword that is a generalization of the training sample.
[0007] In one aspect, obtaining a plurality of training samples includes, for each training sample: providing a seed keyword as an input to a retrieval process; accessing a data repository that stores an association of each of a plurality of queries with one or more retrieval terms; and selecting one or more queries as training samples based on the one or more retrieval terms and the seed keyword.
[0008] In one aspect, training the language model to generate the trained language model includes: for each training sample, training the model to predict the seed keyword provided as an input to the retrieval process to obtain one or more queries as the training sample.
[0009] In one aspect, each query selected from the query repository is a query that has been received as an input query to a search process and generated by a user of the search process.
[0010] In one aspect, determining one or more retrieval terms of a query from the one or more generalizations predicted for the query includes determining that each of the one or more generalizations is a retrieval term.
[0011] In one aspect, determining one or more retrieval terms of a query from the one or more generalizations predicted for the query includes determining one or more retrieval terms from the one or more generalizations, where the one or more retrieval terms define a set of terms different from the set of terms defined by the one or more generalizations.
[0012] In one aspect, for each query, determining one or more retrieval terms of the query from the one or more generalizations predicted for the query includes, for each query: selecting one or more context queries for the query; providing the query and the one or more context queries as inputs to the language model; and receiving from the language model one or more generalizations predicted for the query based on the query and the one or more context queries.
[0013] In one aspect, the language model is one of a multi-task unified model, a zero-shot model, a domain-specific model, or a language representation model.
[0014] Specific embodiments of the subject matter described in this specification can be implemented to achieve one or more of the following advantages. Since the language model learns emerging semantics from the corpus of queries, the techniques discussed in this specification eliminate the need for periodic readjustment and tuning of the relevance-based generalization process. Additionally, by training on the actual queries that have been received and generating retrieval tokens for these queries, the likelihood of hallucination is reduced or even eliminated. Further, once the model is trained, the model is used to process the corpus of queries, thereby generating retrieval tokens for each query. The queries are then associated with their respective retrieval tokens and saved in a data storage system for parallel processing and retrieval. Thus, there is no need to execute the language model at runtime, saving processing resources.
[0015] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a block diagram of an example environment in which the process of generating retrieval tokens can be performed.
[0017] Figure 2A is an illustration of a user interface for a keyword selection process.
[0018] Figure 2B is a system flowchart of a keyword selection process.
[0019] Figure 3 is a system flowchart of a process for training a language model to predict generalizations of queries.
[0020] Figure 4 is a flowchart of a process for generating retrieval tokens using a language model.
[0021] Figure 5 is a block diagram of an example computer.
[0022] Like reference numerals and names in the various drawings indicate like elements. DETAILED DESCRIPTION
[0023] The techniques described in this written specification relate to large language models (LLMs) that are trained to predict generalizations of common queries. The generalizations are then used to generate retrieval tokens for annotating the common queries. By training on the common queries to generate retrieval tokens, the likelihood that the LLM will hallucinate on queries that are in fact invalid is eliminated or greatly reduced. This is because the retrieval tokens generated or derived from the LLM output are used to return the common queries at runtime, rather than the LLM itself returning the predicted query as output.
[0024] Once trained, the LLM is used to generate generalizations for each query in a query set to be annotated. Once retrieval terms are generated from the generalizations, the retrieval terms are used to annotate the queries. The annotated queries are then saved in a data storage system in the search system. At runtime, when the system needs to generate keywords to use based on a seed input, there is no need to use the LLM to generate the keywords because the annotations and queries are accessed through the data repository.
[0025] In an implementation, the LLM is trained on an existing public query corpus and associated retrieval terms. Seed keywords are used to retrieve a subset of the public queries. Specifically, the training samples used to train the LLM are generated by providing the seed keywords to an existing retrieval process. For each seed keyword, the retrieval process selects one or more queries as training samples based on one or more retrieval terms and the seed keyword. The LLM is then trained to predict the seed keyword based on the training samples.
[0026] These features and additional features are described in more detail below.
[0027] As used throughout this document, the phrase "digital component" refers to discrete units of digital content or digital information (e.g., video clips, audio clips, multimedia clips, game content, images, text, points, artificial intelligence output, language model output, or another content unit). Digital components can be stored electronically as a single file or in a collection of files in a physical memory device, and digital components can take the form of video files, audio files, multimedia files, image files, or text files and include advertising information, such that an advertisement is a type of digital component.
[0028] Figure 1 is a block diagram of an example environment 100 in which the process of generating retrieval terms can be performed. The example environment 100 includes a network 102, such as a local area network (LAN), wide area network (WAN), the Internet, or a combination thereof. The network 102 connects an electronic document server 104, a user device 106, a digital component server 108, and a service device 110. The example environment 100 can include many different electronic document servers 104, user devices 106, and digital component servers 108.
[0029] The client device 106 is an electronic device capable of requesting and receiving online resources via the network 102. Example client devices 106 include personal computers, gaming devices, mobile communication devices, digital assistant devices, augmented reality devices, virtual reality devices, and other devices capable of sending and receiving data via the network 102. The client device 106 typically includes a user application (such as a web browser) to facilitate sending and receiving data via the network 102, but native applications (other than the browser) executed by the client device 106 can also facilitate sending and receiving data via the network 102.
[0030] A gaming device is a device that enables a user to participate in a gaming application. For example, in such a device, the user controls one or more characters, avatars, or other rendered content presented in the gaming application. A gaming device typically includes a computer processor, a memory device, and a controller interface (physical or visually rendered) that enables the user to control the content rendered by the gaming application. The gaming device can locally store and execute the gaming application, or execute a gaming application that is at least partially stored and / or served by a cloud server (e.g., an online gaming application). Similarly, the gaming device can interface with a gaming server that executes the gaming application and "streams" the gaming application to the gaming device. The gaming device can be a tablet device, a mobile telecommunications device, a computer, or another device that performs other functions in addition to executing the gaming application.
[0031] A digital assistant device includes a device having a microphone and a speaker. A digital assistant device is typically capable of receiving input by voice, and responding to content using audible feedback, and can present other audible information. In some cases, the digital assistant device also includes a visual display or communicates with a visual display (e.g., via a wireless or wired connection). When a visual display is present, feedback or other information can also be provided visually. In some cases, the digital assistant device can also control other devices, such as lights, locks, cameras, climate control devices, alarm systems, and other devices registered with the digital assistant device.
[0032] As shown, the client device 106 is presenting an electronic document 150. An electronic document is data that presents a set of content at the client device 106. Examples of electronic documents include web pages, word processing documents, Portable Document Format (PDF) documents, images, videos, search result pages, and feeds. Native applications (e.g., "apps" and / or gaming applications) (such as applications installed on mobile, tablet, or desktop computing devices) are also examples of electronic documents. The electronic document can be provided to the client device 106 by an electronic document server 104 ("Electronic Doc Servers").
[0033] For example, the electronic document server 104 may include a server that hosts a publisher's website. In this example, the client device 106 may initiate a request for a given publisher web page, and the electronic server 104 that hosts the given publisher web page may respond to the request by sending machine-executable instructions that initiate the rendering of the given web page at the client device 106.
[0034] In another example, the electronic document server 104 may include an app server from which the client device 106 can download an app. In this example, the client device 106 may download the files required to install the app at the client device 106 and then locally (e.g., on the client device) execute the downloaded app. Alternatively or additionally, the client device 106 may initiate a request to execute the app, and the request is transmitted to a cloud server. In response to receiving the request, the cloud server may execute the app and stream the user interface of the app to the client device 106 so that the client device 106 does not have to execute the app itself. Instead, the client device 106 may present the user interface generated by executing the app on the cloud server and transmit any user interactions with the user interface back to the cloud server for processing.
[0035] An electronic document may include a variety of content. For example, the electronic document 150 may include native content 152 that is located within the electronic document 150 itself and / or does not change over time. The electronic document may also include dynamic content that may change over time or according to each request. For example, the publisher of a given electronic document (e.g., the electronic document 150) may maintain a data source for populating portions of the electronic document. In this example, a given electronic document may include a script, such as the script 154, that causes the client device 106 to request content (e.g., digital components) from the data source when the given electronic document is processed (e.g., rendered or executed) by the client device 106 (or the cloud server). The client device 106 (or the cloud server) integrates the content (e.g., digital components) obtained from the data source into the given electronic document to create a composite electronic document that includes the content obtained from the data source.
[0036] In some cases, a given electronic document (e.g., electronic document 150) may include a digital component script (e.g., script 154) that references the service device 110 or a particular service provided by the service device 110. In these cases, the digital component script is executed by the client device 106 when the given electronic document is processed by the client device 106. Execution of the digital component script configures the client device 106 to generate a request for a digital component 112 (referred to as a "component request"), which is transmitted over the network 102 to the service device 110. For example, the digital component script may enable the client device 106 to generate a packetized data request that includes a header and payload data. The component request 112 may include event data specifying characteristics such as the name (or network location) of the server from which the digital component is being requested, the name (or network location) of the requesting device (e.g., client device 106), and / or information that the service device 110 can use to select one or more digital components or other content to provide in response to the request. The component request 112 is transmitted by the client device 106 over the network 102 (e.g., a telecommunications network) to the server of the service device 110.
[0037] The component request 112 may include event data specifying other event characteristics, such as characteristics of the electronic document being requested and the location within the electronic document where the digital component may be presented. For example, event data specifying a reference (e.g., a URL) to the electronic document (e.g., a web page) in which the digital component will be presented, the available location within the electronic document where the digital component can be presented, the size of the available location, and / or the media type eligible to be presented at the location may be provided to the service device 110. Similarly, event data specifying keywords associated with the electronic document ("document keywords") or entities (e.g., people, places, or things) referenced by the electronic document may also be included in the component request 112 (e.g., as payload data) and provided to the service device 110 to facilitate identification of digital components eligible to be presented with the electronic document. The event data may also include a search query submitted from the client device 106 to obtain a search results page. For example, a query received by a user of the search system may be stored in the query storage device 111.
[0038] The component request 112 may also include event data related to other information, such as information provided by a user of the client device, geographical information indicating the state or region where the component request is submitted, or other information about the context providing the environment in which the digital component will be displayed (e.g., the time of day of the component request, the date of the week of the component request, the type of device on which the digital component will be displayed (such as a mobile device or a tablet device)). The component request 112 may be transmitted, for example, over a packetized network, and the component request 112 itself may be formatted as packetized data having a header and payload data. The header may specify the destination of the packet, and the payload data may include any of the information discussed above.
[0039] The service device 110 selects digital components (e.g., third-party content such as video files, audio files, images, text, game content, augmented reality content, and combinations thereof, all of which may take the form of advertising content or non-advertising content) to be presented with a given electronic document (e.g., at a location specified by the script 154) in response to receiving the component request 112 and / or using the information included in the component request 112.
[0040] In some implementations, the digital components are selected in less than a second to avoid errors that may result from a delayed selection of the digital components. For example, a delay in providing the digital components in response to the component request 112 may cause a page load error at the client device 106 or result in parts of the electronic document remaining unfilled even after other parts of the electronic document are presented at the client device 106.
[0041] Moreover, as the delay in providing the digital components to the client device 106 increases, it is more likely that the electronic document will no longer be presented at the client device 106 by the time the digital components are delivered to the client device 106, thereby negatively affecting the user's experience of the electronic document. Additionally, a delay in providing the digital components may cause the delivery of the digital components to fail, for example, if the electronic document is no longer presented at the client device 106 when the digital components are provided.
[0042] In some implementations, the service device 110 is implemented in a distributed computing system that includes, for example, a server and a collection of multiple computing devices 114 that are interconnected and that identify and distribute digital components in response to the request 112. The collection of multiple computing devices 114 operate together to select from millions of available digital components (DC 1-x) Identify a set of digital components eligible for presentation in an electronic document from a corpus. For example, millions of available digital components can be indexed in a digital component database 116. Each digital component index entry can reference the corresponding digital component and / or include distribution parameters (DP 1 to DP x ) that contribute to (e.g., trigger, condition, or limit) the distribution / transmission of the corresponding digital component. For example, the distribution parameters can contribute to (e.g., trigger) the transmission of a digital component by requiring that a component request include at least one criterion that matches (e.g., exactly or at some pre-specified level of similarity) one of the distribution parameters of the digital component.
[0043] In some implementations, the distribution parameters for a particular digital component can include distribution keywords that must match (e.g., match the electronic document, document keywords, or terms specified in the component request 112) in order for the digital component to be eligible for presentation. Additionally or alternatively, the distribution parameters can include embeddings that can use various different data dimensions, such as website details and / or consumption details (e.g., page viewport, user scroll speed, or other information regarding data consumption). The distribution parameters can also require that the component request 112 include information specifying a particular geographic region (e.g., country or state) and / or information specifying that the component request 112 originated from a particular type of client device (e.g., mobile device or tablet device) in order for the digital component to be eligible for presentation. The distribution parameters can also specify an eligibility value (e.g., a ranking score, or some other specified value) that is used to evaluate the eligibility of a digital component for distribution / transmission (e.g., as well as other available digital components).
[0044] The identification of eligible digital components can be split into multiple tasks 117a to 117c, which are then assigned among the computing devices within a set of multiple computing devices 114. For example, different computing devices 114 in the cluster can each analyze a different portion of the digital component database 116 to identify various digital components having distribution parameters that match the information included in the component request 112. In some implementations, each given computing device 114 in the cluster can analyze a different data dimension (or set of dimensions) and pass (e.g., transmit) the results (Res 1 to Res 3) 118a to 118c of the analysis back to the service device 110. For example, the results 118a to 118c provided by each of the computing devices 114 in the cluster can identify a subset of digital components eligible for distribution in response to the component request and / or a subset of digital components having certain distribution parameters. The identification of the subset of digital components can include, for example, comparing event data with the distribution parameters and identifying a subset of digital components having distribution parameters that match at least some of the characteristics of the event data.
[0045] The service device 110 aggregates the results 118a - 118c received from a set of multiple computing devices 114 and uses the information associated with the aggregated results to select one or more digital components to be provided in response to the request 112. For example, the service device 110 may select a winning set of digital components (one or more digital components) based on the results of one or more content evaluation processes, as described below. Further, the service device 110 may generate and transmit, via the network 102, reply data 120 (e.g., digital data representing the reply), which enables the client device 106 to integrate the winning set of digital components into a given electronic document such that the winning set of digital components (e.g., winning third - party content) and the content of the electronic document are presented together at the display of the client device 106.
[0046] In some implementations, the client device 106 executes the instructions included in the reply data 120, which configure the client device 106 and enable it to obtain the winning set of digital components from one or more digital component servers 108. For example, the instructions in the reply data 120 may include a network location (e.g., a Uniform Resource Locator (URL)) and a script that causes the client device 106 to transmit a server request (SR) 121 to the digital component server 108 to obtain a given winning digital component from the digital component server 108. In response to the request, the digital component server 108 will identify the given winning digital component specified in the server request 121 (e.g., within a database storing multiple digital components) and transmit digital component data (DC data) 122 to the client device 106, and the DC data presents the given winning digital component in the electronic document at the client device 106.
[0047] When the client device 106 receives the digital component data 122, the client device will render the digital component (e.g., third - party content) and present the digital component at the location specified by or assigned to the script 154. For example, the script 154 may create a walled - garden environment (such as a frame) that is presented within the native content 152 of the electronic document 150, e.g., next to the native content. In some implementations, the digital component overlays (or is adjacent to) a portion of the native content 152 of the electronic document 150, and the service device 110 may specify the presentation location within the electronic document 150 in the reply 120. For example, when the native content 152 includes video content, the service device 110 may specify a location or object within the scene depicted in the video content above which the digital component will be presented.
[0048] The service device 110 may also include an artificial intelligence system 160 configured to generate data for identifying digital components. As described in more detail in this specification, the artificial intelligence (“AI”) system 160 may generate retrieval tokens for a query and annotate the query. The provider of the digital components then uses these annotations to identify keywords for identifying a particular digital component. The keywords may be a single word or term, or a combination of a term and a number.
[0049] A large language model (“LLM”) is a model trained to generate and understand human language. LLMs are trained on large datasets of text and code, and they can be used for a variety of tasks. For example, an LLM can be trained to translate text from one language to another; summarize text, such as website content, search results, news articles, or research papers; answer questions about text, such as “What is the capital of Georgia?”; create chatbots that can converse with humans; and generate creative text, such as poems, stories, and code. LLMs can also be trained for specific tasks. In the implementations described in this document, the LLM is trained to predict a generalization of the input text. For example, for an input 172 of one or more queries, the LLM may generate a generalization of the input 172 as an output 174.
[0050] The language model 170 may be any suitable language model neural network that receives an input sequence composed of text tokens selected from a vocabulary and autoregressively generates an output sequence composed of text tokens from the vocabulary. For example, the language model 170 may be a Transformer-based language model neural network or a recurrent neural network-based language model.
[0051] In some cases, when the neural network used to implement the language model 170 autoregressively generates an output sequence of tokens, the language model 170 may be referred to as an autoregressive neural network. More specifically, the autoregressive-generated output is created by generating each particular token in the output sequence conditioned on the current input sequence, which includes any tokens in the output sequence that are before the particular token (i.e., tokens that have been generated for any previous positions in the output sequence before the particular position of the particular token), and a context input that provides context for the output sequence.
[0052] For example, when generating a token at any given position in the output sequence, the current input sequence can include tokens at any previous position in the input sequence and the output sequence that are before the given position. As a specific example, the current input sequence can include the input sequence followed by tokens at any previous position in the output sequence that are before the given position. Optionally, the input and the current output sequence can be separated by one or more predetermined tokens within the current input sequence.
[0053] More specifically, to generate a particular token at a particular position within the output sequence, the neural network of the language model 170 can process the current input sequence to generate a score distribution (e.g., a probability distribution) that assigns a corresponding score (e.g., a corresponding probability) to each token in the token vocabulary. Then, the neural network of the language model 170 can use the score distribution to select a token from the vocabulary as the particular token. For example, the neural network of the language model 170 can greedily select the token with the highest score, or can sample tokens from the distribution using, for example, nucleus sampling or another sampling technique.
[0054] As a specific example, the language model 170 can be a neural network based on an autoregressive Transformer that includes (i) a plurality of attention blocks, each attention block applying self-attention operations; and (ii) an output subnetwork that processes the output of the last attention block to generate a score distribution.
[0055] The language model 170 can have any of a variety of Transformer-based neural network architectures. Examples of such architectures include those described in: J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark et al., "Training compute-optimal large language models", arXiv preprint arXiv:2203.15556, 2022; J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, H. F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche, L. A. Hendricks, M. Rauh, P. Huang, A. Glaese, J. Welbl, S. Dathathri, S. Huang, J. Uesato, J. Mellor, I. Higgins, A. Creswell, N. McAleese, A. Wu, E. Elsen, S. M. Jayakumar, E. Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M. Paganini, L. Sifre, L. Martens, X. L. Li, A. Kuncoro, A. Nematzadeh, E. Gribovskaya, D. Donato, A. Lazaridou, A. Mensch, J. Lespiau, M. Tsimpoukelli, N. Grigorev, D. Fritz, T. Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D. Toyama, C. de Masson d’Autume, Y. Li, T. Terzi, V. Mikulik, I. Babuschkin, A. Clark, D. de Las Casas, A. Guy, C. Jones, J. Bradbury, M.Johnson, B. A., Hechtman, L., Weidinger, I., Gabriel, W. S., Isaac, E., Lockhart, S., Osindero, L., Rimell, C., Dyer, C., Vinyals, O., Ayoub, K., Stanway, J., Bennett, L., Hassabis, D., Kavukcuoglu, K., & Irving, G. (2021). Scaling language models: Methods, analysis & insights from training gopher. CoRR, abs / 2112.11446; Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., & Liu, P. J. (2019). Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683; Adiwardana, D., Luong, M.-T., So, D. R., Hall, J., Fiedel, N., Thoppilan, R., Yang, Z., Kulshreshtha, A., Nemade, G., Lu, Y., & Le, Q. V. (2020). Towards a human-like open-domain chatbot. CoRR, abs / 2001.09977, 2020; and Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al., "Language models are few-shot learners", arXiv preprint arXiv:2005.14165, 2020. Such example models include multi-task unified models, zero-shot models, domain-specific models, or language representation models.
[0056] However, generally, a Transformer-based neural network includes a sequence of attention blocks, and during processing of a given input sequence, each attention block in the sequence receives corresponding input hidden states of each input token in the given input sequence. The attention block then updates each hidden state in the hidden states at least in part by applying self-attention to generate corresponding output hidden states for each input token in the input tokens. The input hidden state of the first attention block is the embedding of the input tokens in the input sequence, while the input hidden state of each subsequent attention block is the output hidden state generated by the previous attention block.
[0057] In this example, the output subnetwork processes the output hidden state generated by the last attention block in the sequence for the last input token in the input sequence to generate a score distribution.
[0058] Typically, since the language model is autoregressive, the service device 110 can use the same language model 170 to generate multiple different candidate output sequences in response to the same request, e.g., by using beam search decoding from the score distribution generated by the language model 170, using sampling and ranking decoding strategies, using different random seeds for the pseudo-random number generator used when sampling different runs of the language model 170, or using another decoding strategy that exploits the autoregressive nature of the language model.
[0059] In some implementations, the language model 170 is pre-trained, i.e., trained on a language modeling task that does not require providing evidence in response to a user question, and the service device 110 (e.g., using the AI system 160) causes the language model 170 to generate an output sequence according to a predetermined grammar through natural language prompts in the input sequence.
[0060] For example, a service device 110 (e.g., an AI system 160) or a separate training system pre-trains a language model 170 (e.g., a neural network) on a language modeling task, such as a task that requires predicting the next token in the training data after the current sequence given the current sequence of text tokens. As a specific example, the language model 170 can be pre-trained on a maximum likelihood objective on a large text dataset (e.g., text publicly available from the Internet or another text corpus).
[0061] As described above, the language model 170 can generate retrieval tokens for a query and annotate the query. Then, providers of digital components use these annotations to identify keywords for identifying specific digital components. One way for a provider to identify keywords is to use a keyword selection process. Refer to Figure 2A and Figure 2B for a description of an example keyword selection process. Specifically, Figure 2A is an illustration of a user interface 200 for the keyword selection process, while Figure 2B is a system flowchart 250 of the keyword selection process.
[0062] In Figure 2A , the user has entered the seed "leather boots" into the keyword selection process user interface 200. The user interface 200 includes three columns - a keyword column 202, an average monthly search column 204, and a six-month change column 206. The keyword column 202 lists the keywords returned in response to the seed "leather boots". In some implementations, the keywords are common queries that have been received from users. The average monthly search column 204 and the six-month change column 206 display statistics related to a particular keyword. The user can evaluate the statistics and the keywords and determine whether to select the keyword to include in the criteria for selecting digital components. For example, if the digital component is an advertisement, the selected keyword can be associated with an advertising campaign that includes the advertisement.
[0063] Keywords can be selected by checking the corresponding selection box 208. Additional keywords can be displayed by selecting the "Show More" button 210.
[0064] In Figure 2BIn the system flow chart, the keyword planner process 252 submits seeds to the keyword retrieval server 254. Here, the seed is "women’s jewelry". The keyword retrieval server 254 accesses the data storage device 256 that stores the annotated queries. The queries are annotated in the form of {Q:RT1, RT2…RTn}, where RT is a retrieval term. As used herein, a retrieval term is a term used to annotate a query, such as a single word, a combination of words and terms, or even a phrase. The retrieval terms of a particular query can be interpreted as semantic generalizations of that query. A particular retrieval term can correspond to a particular word in a multi-word query, or can correspond to two or more words in a multi-word query.
[0065] The keyword retrieval server 254 can use one or more evaluation processes to select keywords based on the seeds. These evaluation processes can include a relevance evaluation process, a word matching process, a query quality evaluation process, and other evaluation processes that can be used to evaluate the response to the seeds. As Figure 2B shown, the seed "women's jewelry" returned the keywords "ladies gold ring", "ladies gold necklace", and "women’s watch". Although only three keywords are shown, in reality, dozens, hundreds, or even thousands of keywords can be returned in response to one seed.
[0066] Figure 3 is a system flow chart of a process 300 for training a language model to predict generalizations of queries. Refer Figure 4 to describe process 300, Figure 4 is a flow chart of a process 400 for generating retrieval terms using a language model. Process 400 is implemented in a computer system that typically includes multiple processors.
[0067] In operation, process 400 obtains training samples (402). Each training sample includes one or more queries, and each query in the query is selected from a query repository (such as, Figure 1 the query storage device 111). Each query is a query that has been received as an input query to a search process and is generated by a user of the search process.
[0068] An example method of generating training samples is to provide seeds to the keyword planner process (e.g., Figure 2B process 250). For example, each training sample is generated by providing seed keywords to the retrieval process. As Figure 3As shown, for the seed "women's military boots", training samples {Brand X Shape 91s; women's black combat boots; women's summer boots UK} are generated. The term "Brand X" is a brand name, and the term "Shape 91s" is a product name. Training samples are generated by accessing a query data repository 256 annotated with one or more search tokens each, and queries are selected based on one or more search tokens and seed keywords. Training samples 302 are generated for multiple seeds in order to generate sufficient training samples to train the language model 170.
[0069] Process 400 trains the language model 170 to predict one or more generalizations (404) that describe the training samples for each training sample. Each generalization is an n-gram of one or more terms. For example, as Figure 3 shown, for the training sample {BrandX Shape 91s; women's black combat boots; women's summer boots UK}, the generated generalization 306 is {women's shoes; women's boots}. Figure 3 The generalizations shown are for the model 170 that is about to complete the training process. However, when the training process begins, the generalizations may be far less accurate.
[0070] In some implementations, the model 170 is trained such that one or more generalizations include the seed keyword used to generate the training sample. For example, for the training sample {Brand X Shape91s; women's black combat boots; women's summer boots UK} corresponding to the seed keyword "women's military boots", the output of the model 170 will be one or more generalizations that include the seed keyword "women's military boots" as a generalization.
[0071] In some implementations, the language model 170 is trained to predict generalizations and then the generalizations are post-processed to determine search tokens. For example, the generalizations can be processed to remove stop words such as "the", "and", etc., and the generalizations can be processed to remove overly general words or terms such as "thing", "product", etc.
[0072] In other implementations, the language model 170 can be trained to predict generalizations that do not require post-processing to determine search tokens. Instead, the prediction makes each of one or more generalizations be a search token.
[0073] Once the model 170 is trained, queries stored in the query repository are processed through the trained model 170 to predict one or more generalizations (408) for each query. For example, a single query can be input into the trained model 170 and the model can predict a generalization based on the single query.
[0074] In another implementation, the query to be annotated may include one or more other queries selected as context queries. Then, the query and the one or more context queries are input into the language model 170, and one or more generalizations predicted for the query based on the query and the one or more context queries are received.
[0075] Then, for each query, process 400 determines one or more retrieval terms (408) for the query from the one or more generalizations predicted for the query. As described above, the language model 170 can be trained to predict generalizations, and then the generalizations are post-processed to determine retrieval terms. For example, the generalizations can be processed to remove stop words such as "the", "and", etc., and the generalizations can be processed to remove overly general words or terms such as "thing", "product", etc. In other words, the one or more retrieval terms define a set of terms different from the set of terms defined by the one or more generalizations.
[0076] Alternatively, the language model 170 can be trained to predict generalizations that do not require post-processing to determine retrieval terms. Instead, the prediction is such that each of the one or more generalizations is a retrieval term. In other words, the one or more retrieval terms define a set of terms that is the same as the set of terms defined by the one or more generalizations.
[0077] Then, for each query, process 400 stores an association of the query with the one or more retrieval terms determined for the query in the data repository (410). For example, process 400 can store the annotated query in Figure 2B data repository 256. Once stored, the annotated query can be used for keyword selection, such as as described in Figure 2A and Figure 2B described.
[0078] Figure 5 FIG. is a block diagram of an example computer system 500 that can be used to perform the operations described above. System 500 includes a processor 510, a memory 520, a storage device 530, and an input / output device 540. Each of the components 510, 520, 530, and 540 can be interconnected, for example, using a system bus 550. The processor 510 is capable of processing instructions for execution within the system 500. In one implementation, the processor 510 is a single-threaded processor. In another implementation, the processor 510 is a multi-threaded processor. The processor 510 is capable of processing instructions stored in the memory 520 or on the storage device 530.
[0079] Memory 520 stores information within system 500. In one implementation, memory 520 is a computer-readable medium. In one implementation, memory 520 is a volatile memory unit. In another implementation, memory 520 is a non-volatile memory unit.
[0080] Storage device 530 can provide mass storage for system 500. In one implementation, storage device 530 is a computer-readable medium. In various different implementations, storage device 530 can include, for example, a hard disk device, an optical disk device, a storage device shared by multiple computing devices over a network (e.g., a cloud storage device), or some other mass storage device.
[0081] Input / output device 540 provides input / output operations for system 500. In one implementation, input / output device 540 can include one or more of a network interface device (e.g., an Ethernet card), a serial communication device (e.g., an RS-232 port), and / or a wireless interface device (e.g., an 802.11 card). In another implementation, the input / output device can include a drive device configured to receive input data and send output data to other devices (e.g., a keyboard, a printer, a display, and other peripheral devices 560). However, other implementations can also be used, such as mobile computing devices, mobile communication devices, set-top box TV client devices, etc.
[0082] Although an example processing system is described Figure 5 herein, implementations of the subject matter and functional operations described in this specification can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them.
[0083] An electronic document (referred to herein simply as a document) need not correspond to a file. A document can be stored as part of a file that holds other documents, in a single file dedicated to the document in question, or in multiple coordinated files.
[0084] Regarding the collection and / or use of personal information about users by the systems discussed herein, users may be provided with the opportunity to enable / disable or control programs or features that can collect and / or use personal information (e.g., information about a user's social network, social actions or activities, a user's preferences, or a user's current location). Additionally, certain data can be disposed of in one or more ways before it is stored or used so that personally identifiable information associated with the user is deleted. For example, a user's identity can be anonymized so that the user's personally identifiable information cannot be determined, or, if location information is obtained, the user's geographical location can be generalized (such as to the city, postal code, or state level) so that the user's specific location cannot be determined.
[0085] Embodiments of the subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents), or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on a computer storage medium for execution by, or to control the operation of, a data processing apparatus. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to an appropriate receiver apparatus for execution by a data processing apparatus. A computer storage medium can be, or can include in whole or in part, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, although a computer storage medium is not a propagated signal, a computer storage medium can be the source or destination of computer program instructions encoded in an artificially generated propagated signal. A computer storage medium can also be, or can include in whole or in part, one or more separate physical components or media (such as multiple CDs, disks, or other storage devices).
[0086] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0087] The term "data processing apparatus" includes all types of apparatus, devices, and machines for processing data, including, by way of example, programmable processors, computers, system-on-a-chip, or multiple ones or combinations of the foregoing. The apparatus may include dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). In addition to hardware, the apparatus may also include code that creates an execution environment for the computer programs involved, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and the execution environment may implement various different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.
[0088] This document relates to service apparatuses. As used herein, a service apparatus is one or more data processing apparatuses that perform operations to facilitate the distribution of content over a network. The service apparatus is depicted as a single block in the block diagram. However, although the service apparatus may be a single device or a single set of devices, the present disclosure contemplates that the service apparatus may also be a group of devices, or even multiple different systems that communicate to provide various content to client devices. For example, the service apparatus may encompass one or more of a search system, a video streaming service, an audio streaming service, an email service, a navigation service, an advertising service, a gaming service, or any other service.
[0089] A computer program (also called a program, software, software application, script, or code) can be written in any form of programming language (including compiled or interpreted languages or declarative or procedural languages); and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may or may not correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers distributed at one site or across multiple sites and interconnected by a communication network.
[0090] The processes and logical flows described in this specification can be performed by one or more programmable processors that execute one or more computer programs to perform actions by operating on input data and generating output. The processes and logical flows can also be performed by dedicated logic circuitry, and the apparatus can also be implemented as dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0091] For example, processors suitable for executing computer programs include general and special purpose microprocessors, as well as any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The basic elements of a computer are a processor for performing operations in accordance with the instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data therefrom or to transfer data thereto or both. However, a computer need not have such devices. In addition, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0092] To provide for interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user may provide input to the computer. Other kinds of devices may also be used to provide for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input received from the user may be in any form, including sound, speech, or tactile input. Additionally, a computer may interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending web pages to a web browser on a client device of the user in response to a request received from the web browser.
[0093] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes backend components, such as a data server, or includes middleware components, such as an application server, or includes frontend components, such as a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0094] The computing system can include clients and servers. Clients and servers are typically located remotely from each other and typically interact through a communication network. The relationship between a client and a server arises from computer programs that run on respective computers and have a client-server relationship with each other. In some embodiments, the server transmits data (e.g., an HTML web page) to a client device (e.g., for the purpose of displaying the data to a user interacting with the client device and receiving user input from the user). Data generated at the client device (e.g., the result of a user interaction) can be received at the server from the client device.
[0095] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination can be excluded from the combination, and the claimed combination can cover a sub-combination or a variation of a sub-combination.
[0096] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order, or that all illustrated operations be performed, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system components in the embodiments described above should not be understood as required in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0097] Accordingly, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result. In addition, the processes depicted in the figures need not be in the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing may be advantageous.
Claims
1. A computer-implemented method comprising: obtaining a plurality of training samples, each training sample comprising one or more queries, each query of the one or more queries being a query selected from a query repository; Training a language model to generate a trained language model, wherein the trained language model predicts, for each training sample, one or more generalizations describing the training sample, each of the one or more generalizations being an n-gram; processing the queries stored in the query repository through the trained language model to predict, for each query, one or more generalizations of the query; For each query, determining one or more search terms for the query from the one or more generalizations predicted for the query; as well as For each query, an association of the query with the one or more search terms determined for the query is stored in the data repository.
2. The computer-implemented method of claim 1 , wherein training the language model to generate the trained language model comprises: For each training sample, the model is trained to predict a seed keyword that is a generalization of the training sample.
3. The computer-implemented method of claim 1 , wherein obtaining a plurality of training samples comprises, for each training sample: providing seed keywords as input to the search process; accessing a data repository that stores, for each query in a plurality of queries, an association of the query with the one or more search terms; and One or more queries are selected as training samples based on the one or more search terms and the seed keyword.
4. The computer-implemented method of claim 3, wherein training the language model to generate the trained language model comprises, for each training sample: The model is trained to predict the seed keywords provided as input to the retrieval process to obtain the one or more queries as the training samples.
5. The computer-implemented method of claim 1, wherein each query selected from the query repository is a query that has been received as an input query to a search process and generated by a user of the search process.
6. The computer-implemented method of claim 1, wherein determining one or more search terms for the query from the one or more generalizations predicted for the query comprises determining that each of the one or more generalizations is a search term.
7. A computer-implemented method as described in claim 1, wherein determining one or more search terms for the query from the one or more generalizations predicted for the query includes determining one or more search terms from the one or more generalizations, wherein the one or more search terms define a set of terms that is different from the set of terms defined by the one or more generalizations.
8. The computer-implemented method of claim 1 , wherein for each query, determining one or more search terms for the query from the one or more generalizations predicted for the query comprises: selecting one or more context queries for the query; providing the query and the one or more context queries as input to the language model; as well as One or more generalizations predicted for the query based on the query and the one or more context queries are received from the language model.
9. The computer-implemented method of claim 1, wherein the language model is one of a multi-task unified model, a zero-shot model, a domain-specific model, or a language representation model.
10. A system comprising: one or more computers, the one or more computers being in data communication; as well as One or more non-transitory computer-readable media storing instructions executable by the one or more computers, the instructions, when executed by the one or more computers, causing the one or more computers to perform operations comprising: obtaining a plurality of training samples, each training sample comprising one or more queries, each query of the one or more queries being a query selected from a query repository; Training a language model to generate a trained language model, wherein the trained language model predicts, for each training sample, one or more generalizations describing the training sample, each of the one or more generalizations being an n-gram; processing the queries stored in the query repository through the trained language model to predict, for each query, one or more generalizations of the query; For each query, determining one or more search terms for the query from the one or more generalizations predicted for the query; and For each query, an association of the query with the one or more search terms determined for the query is stored in the data repository.
11. The one or more non-transitory computer-readable media of claim 10, wherein training the language model to generate the trained language model comprises: For each training sample, the model is trained to predict a seed keyword that is a generalization of the training sample.
12. The one or more non-transitory computer-readable media of claim 10, wherein obtaining a plurality of training samples comprises, for each training sample: providing seed keywords as input to the search process; accessing a data repository that stores, for each query in a plurality of queries, an association of the query with the one or more search terms; and One or more queries are selected as training samples based on the one or more search terms and the seed keyword.
13. The one or more non-transitory computer-readable media of claim 12, wherein training the language model to generate the trained language model comprises, for each training sample: The model is trained to predict the seed keywords provided as input to the retrieval process to obtain the one or more queries as the training samples.
14. The one or more non-transitory computer-readable media of claim 10, wherein each query selected from the query repository is a query that has been received as an input query to a search process and generated by a user of the search process.
15. The one or more non-transitory computer-readable media of claim 10, wherein determining one or more search terms for the query from the one or more generalizations predicted for the query comprises determining that each of the one or more generalizations is a search term.
16. One or more non-transitory computer-readable media as described in claim 10, wherein determining one or more search terms for the query from the one or more generalizations predicted for the query includes determining one or more search terms from the one or more generalizations, wherein the one or more search terms define a set of terms that is different from the set of terms defined by the one or more generalizations.
17. The one or more non-transitory computer-readable media of claim 10, wherein for each query, determining one or more search terms for the query from the one or more generalizations predicted for the query comprises: selecting one or more context queries for the query; providing the query and the one or more context queries as input to the language model; as well as One or more generalizations predicted for the query based on the query and the one or more context queries are received from the language model.
18. The one or more non-transitory computer-readable media of claim 10, wherein the language model is one of a multi-task unified model, a zero-shot model, a domain-specific model, or a language representation model.
19. One or more non-transitory computer-readable media storing instructions that, when executed by an artificial intelligence system, cause the artificial intelligence system to perform operations comprising: obtaining a plurality of training samples, each training sample comprising one or more queries, each query of the one or more queries being a query selected from a query repository; Training a language model to generate a trained language model, wherein the trained language model predicts, for each training sample, one or more generalizations describing the training sample, each of the one or more generalizations being an n-gram; processing the queries stored in the query repository through the trained language model to predict, for each query, one or more generalizations of the query; For each query, determining one or more search terms for the query from the one or more generalizations predicted for the query; as well as For each query, an association of the query with the one or more search terms determined for the query is stored in the data repository.