Specific perception teacher model and student model based on large language model

Generating a specific perceived teacher model through large language models solves the problem that existing models are difficult to distinguish between query and login page specificity, realizes a faster and less resource-based training process, generates query-to-document scores with specific tags, and improves prediction performance.

CN120476405APending Publication Date: 2025-08-12GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380065665.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-08-17
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Existing models are difficult to accurately distinguish the specificity of query and the specificity of login pages, resulting in excessive distribution of digital components and inefficient system.

Method used

A large language model is used to generate a specially-aware teacher model. By training the teacher model and student model, using the specific deviation of query and login pages, a specific marker query to document score is generated, reducing resource and time requirements are reduced, and a larger training set is built to improve prediction performance.

Benefits of technology

A faster and less resource-based training process is implemented, and a student training set with specific tags query to document scores is generated, improving the performance of predictions for login page queries, and overcoming the inaccuracy of previous student models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120476405A_ABST
    Figure CN120476405A_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating a teacher training set of specifically tagged query-to-document scores for training a teacher model using a large language model (LLM). The teacher model is in turn used to generate a student training set of specifically tagged query document scores for training student models. The student model is integrally trained to predict performance of the query for the login page using the student model-specific training set and at least one other student model training set that does not quantify the query to the login page-specific deviation.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] This specification relates to data processing, and in particular, the use of large language models and refinement to generate specificity-aware student models.

[0002] Digital components, which are discrete units of digital content or digital information, are selected for inclusion in the digital content delivered to the requesting user device. To select the digital components to deliver, the system needs to be able to evaluate which digital components are most suitable for delivery for a particular delivery instance. One method of selecting digital components is by querying landing pages associated with the digital components using a matching query.

[0003] The system can predict whether a query is "relevant" to a landing page based on various modeling techniques. As used herein, determining whether a query is relevant to a landing page is based in part on the likelihood that the landing page satisfies the user information need of the user who issued the query to the information processing system. This determination can be based on traffic logs, semantic analysis of the query and landing page, and other factors.

[0004] However, for a given query and landing page pair determined to be relevant, the specificity of the landing page and the specificity of the query can have a significant impact on the performance of the query and landing page pair. For example, the query "recyclable bags" can refer to many types of bags (grocery bags, garbage bags, resealable bags, etc.) and is therefore more general than a specific landing page. On the other hand, the query "clear recyclable resealable bags" has roughly similar specificity to many specific landing pages, while the query "4×4 clear recyclable resealable plastic bags" is generally more specific than a specific landing page.

[0005] However, existing models struggle to predict whether a query is more general, more specific, or has the same specificity as a landing page. Because existing models can't distinguish between specificity, they tend to favor matches of general queries to specific landing pages. This leads to over-dispatching of numerical components, which in turn makes the system inefficient. Summary of the Invention

[0006] Generally speaking, one innovative aspect of the subject matter described herein can be embodied in a method comprising the following acts: for each of a plurality of query and landing page pairs, each query and landing page pair being a specific query and a specific landing page selected for the query: obtaining, for the query, a proper subset of search results responsive to the query, the proper subset of search results being search results determined to be ranked higher than other search results in a search result set responsive to the query, wherein each search result is linked to a search result page; for each search result page and landing page: generating, by a large language model trained to determine label probabilities, a plurality of label probabilities, each label probability being generated for a specific label in the label set and being generated by a separate inference call of the large language model, Determining a document-to-document score from the plurality of label probabilities, the document-to-document score being a quantification of the specificity of the landing page relative to the specificity of the search result page, and determining a query-to-document score from the document-to-document score, the query-to-document score quantifying the query-to-landing page specificity drift; storing training tuples in a teacher model specificity training set, each training tuple being a query and landing page pair and the query-to-document score determined for the query and landing page pair; and training a teacher model on the teacher model training set to determine a query-to-document score for a query input and a landing page selected for the query, wherein the query-to-document score is determined without search results and search result pages determined to be responsive to the query. The above features may also be embodied in a system and a non-transitory computer storage medium.

[0007] These and other embodiments can each optionally include one or more of the following aspects. In one aspect, the method includes: using the teacher model after training to process a plurality of query and landing page pairs to determine a query-to-document score for each query and landing page pair; and storing training tuples as a student model-specific training set, each training tuple being a query and landing page pair and the query-to-document score determined for the query and landing page pair.

[0008] In another aspect, the method includes ensembling student models to predict performance for queries targeting landing pages using the student model-specific training set and at least one other student model training set that does not quantify the query-to-landing page specific deviation.

[0009] On the other hand, performing the ensemble training on the student model includes performing ensemble training on a dual encoder.

[0010] In another aspect, the dual encoder includes a query encoder that generates a query embedding from a combined embedding of two or more query dimensions; and a document encoder that generates a document embedding from a combined embedding of two or more document dimensions.

[0011] In another aspect, the teacher model includes a crisscross attention model comprising: a transformer layer that receives as input the combined embeddings from two or more query dimensions and two or more document embeddings for an input query and landing page pair; and a pooler layer that receives the output of the transformer layer and generates query-to-document scoring data for the input query and landing page pair.

[0012] On the other hand, the other student model training set includes data labeled with information quality (IQ) and landing page quality (LQ), and the integrated training includes: when the data labeled with information quality (IQ) and the data labeled with landing page quality (LQ) both indicate scores that exceed the corresponding high score threshold, increasing the training loss of the sample; and when the data labeled with information quality (IQ) and the data labeled with landing page quality (LQ) both indicate scores that do not meet the corresponding low score threshold, reducing the training loss of the sample.

[0013] In one aspect, determining the query-to-document score that quantifies query-to-landing page specific deviation from the document-to-document score includes determining a central tendency value for the document-to-document score determined for the query and landing page pair.

[0014] Another innovative aspect of the subject matter described in this specification can be embodied in a method comprising the following acts: for each of a plurality of query and landing page pairs, each query and landing page pair being a particular query and a particular landing page selected for the query: obtaining, for the query, a proper subset of search results responsive to the query, the proper subset of search results being search results determined to be ranked higher than other search results in a search result set responsive to the query, wherein each search result is linked to a search result page; and for each search result page and landing page: generating, by a large language model trained to determine a plurality of label probabilities, a plurality of label probabilities, each label probability being generated for a particular label in the label set and generated by a single inference call of the large language model. The method further comprises: determining a document-to-document score from the plurality of label probabilities, the document-to-document score being a quantification of the specificity of the landing page relative to the specificity of the search result page, determining a query-to-document score from the document-to-document score, the query-to-document score quantifying the query-to-landing page specificity deviation, and storing training tuples as a teacher model specificity training set, each training tuple being a query and landing page pair and the query-to-document score determined for the query and landing page pair; and training a teacher model on the teacher model training set to determine a query-to-document score for a query input and a landing page selected for the query, wherein the query-to-document score is determined without the search results and search result pages determined to be responsive to the query. The above features may also be embodied in a system and a non-transitory computer storage medium.

[0015] These and other embodiments can each optionally include one or more of the following aspects: In one aspect, the large language model is an encoder-only model.

[0016] In another aspect, the method includes: using the teacher model to process a plurality of query and landing page pairs after the teacher model is trained to determine a query-to-document score for each query and landing page pair; and storing the training tuples as a student model-specific training set, each training tuple being a query and landing page pair and the query-to-document score determined for the query and landing page pair.

[0017] In another aspect, the method includes ensembling student models to predict performance for queries targeting landing pages using the student model-specific training set and at least one other student model training set that does not quantify the query-to-landing page specific deviation.

[0018] On the other hand, ensemble training the student model includes ensemble training the dual encoders.

[0019] In another aspect, a dual encoder includes a query encoder that generates a query embedding from a combined embedding of two or more query dimensions; and a document encoder that generates a document embedding from a combined embedding of two or more document embeddings.

[0020] On the other hand, the query dimension includes knowledge graph entities and one or more of query text data, location data, and term data of significant terms, and the document dimension includes knowledge graph entities and one or more of uniform resource locator data, term data of significant terms, and document title data.

[0021] On the other hand, the other student model training set includes data labeled with information quality (IQ) and landing page quality (LQ), and the integrated training includes: when the data labeled with information quality (IQ) and the data labeled with landing page quality (LQ) both indicate scores that exceed the corresponding high score threshold, increasing the training loss of the sample; and when the data labeled with information quality (IQ) and the data labeled with landing page quality (LQ) both indicate scores that do not meet the corresponding low score threshold, reducing the training loss of the sample.

[0022] In another aspect, the teacher model includes a crisscross attention model comprising: a transformer layer that receives as input the combined embeddings from two or more query dimensions and two or more document embeddings for an input query and landing page pair; and a pooler layer that receives the output of the transformer layer and generates query-to-document scoring data for the input query and landing page pair.

[0023] On the other hand, the query dimension includes knowledge graph entities and one or more of query text data, location data, and term data of significant terms; and the document dimension includes knowledge graph entities and one or more of uniform resource locator data, term data of significant terms, and document title data.

[0024] In another aspect, determining the query-to-document score that quantifies query-to-landing page specific deviation from the document-to-document score includes determining a central tendency value for the document-to-document score determined for the query and landing page pair.

[0025] Particular embodiments of the subject matter described in this specification can be implemented to achieve one or more of the following advantages. Training of a teacher model produces a teacher model that can generate a student training set of query-to-document scores labeled with specificity, and the teacher model requires fewer resources and has faster processing time than when a document-to-document large language model is used. This is because the document-to-document large language model predicts document-to-document specificity scores for document pairs individually and then uses these values to generate corresponding query-to-document specificity scores. The teacher model, on the other hand, predicts query-to-document scores based solely on the landing page and the query. Assuming that each landing page has M documents that are scored for specificity, and assuming that each landing page requires at least N specificity inferences, the teacher model produces at least (N×M – M) fewer inferences than when a document-to-document large language model is used to generate a student training set of query-to-document scores labeled with specificity. Thus, when constructing a training dataset, the systems and methods described herein achieve processing optimizations that reduce the time to construct the dataset and reduce the computer resources required to construct the dataset.

[0026] Furthermore, since the teacher model does not require the pages linked to by the search results, it does not require the use of search engine resources when building a larger training set. This bypasses the need for search processing, freeing up computer resources and system bandwidth.

[0027] In some implementations, a document-to-document large language model can be implemented using only an encoder model, and M document-to-document predictions can be determined from a single encoding process rather than M separate inferences. This reduces the processing time required to generate a teacher training set of specifically labeled query-to-document scores for training the teacher model.

[0028] The use of a teacher model also enables the practical generation of a very large student training set of uniquely labeled query-to-document scores. This large training set ensures that the resulting dual-encoder student model, which predicts performance for queries targeting landing pages, is uniquely unique. This uniqueness overcomes the inaccuracies of previous student models, as ensemble training does not account for variations in query and landing page uniqueness.

[0029] In some implementations, the student and teacher models consider the knowledge graph entities of the query and landing page during training. Furthermore, by considering the average number of entities present for a given query and the average number of entities present for a given landing page, the feature lengths of the query and document entities can be selected to avoid training and inference time degradation while still achieving overall performance improvements. This makes the student model's performance more robust than a student model that does not consider knowledge graph entities.

[0030] The overall workflow described in this specification is a staged workflow in which a teacher model that determines query-to-document specificity scores is trained from a training corpus generated from a labeled training corpus of document-to-document scores generated by a large language model. The teacher model is then used to generate a larger corpus of examples of specifically labeled query-to-document scores, which is then used as part of the ensemble training of the student model. Each training stage in the workflow results in a model that generates output using fewer processing resources than its predecessor.

[0031] The details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is a block diagram of an example environment in which a large language model system is used in a training workflow to generate a student model that predicts specificity-aware performance for queries targeting landing pages.

[0033] Figure 2 is a diagram of the workflow for generating specificity-aware models.

[0034] Figure 3 is a diagram of the workflow for determining a document-to-document score for two documents.

[0035] Figure 4 is a diagram of a workflow for determining query-to-document scores for queries and associated landing pages.

[0036] Figure 5 is a flowchart of an example process for training a teacher model that determines a query-to-document score that is partially specificity-aware for an input of a query and a landing page selected for the query.

[0037] Figure 6 is a flowchart of an example process for training a student model that determines a query-to-document score that is partially specificity-aware for an input of a query and a landing page selected for the query.

[0038] Figure 7 is an example architecture for the teacher model.

[0039] Figure 8 is a sample architecture for the Student model.

[0040] Figure 9 is a block diagram of an example computer system.

[0041] Like reference numbers and designations throughout the various drawings indicate like elements. DETAILED DESCRIPTION

[0042] Overview

[0043] The technology described in this written specification involves a large language model (LLM) that is used to generate a teacher training set of uniquely labeled query-to-document scores for training a teacher model. The teacher model is trained to take queries and landing pages as input and generate query-to-document scores based on the input. The scores are then used to label query and landing page pairs to generate a uniquely labeled student training set. The student training set is then used to train an ensemble of student models to predict performance for queries against landing pages.

[0044] As described below, workflows and models are selected to reduce processing resources in a staged manner. Furthermore, in some implementations, certain model architectures are selected to reduce the overall inference time required to generate a training set. Each of these features, whether used individually or in combination, achieves significant savings in both physical computer resource requirements and processing time.

[0045] As used throughout this document, the phrase "digital component" refers to a discrete unit of digital content or digital information (e.g., a video clip, an audio clip, a multimedia clip, game content, an image, text, a gist, an artificial intelligence output, a language model output, or another unit of content). A digital component may be electronically stored in a physical memory device as a single file or in a collection of files, and a digital component may take the form of a video file, an audio file, a multimedia file, an image file, or a text file and include advertising information, such that an advertisement is a type of digital component.

[0046] Sample operating environment

[0047] Figure 1 1 is a block diagram of an example environment in which a large language model system 170 is used in a training workflow to generate a student model that predicts specificity-aware performance for queries targeting landing pages. Example environment 100 includes a network 102, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof. Network 102 connects electronic document servers 104, user devices 106, digital component servers 108, and service devices 110. Example environment 100 may include many different electronic document servers 104, user devices 106, and digital component servers 108.

[0048] Client devices 106 are electronic devices capable of requesting and receiving online resources over network 102. Example client devices 106 include personal computers, gaming devices, mobile communication devices, digital assistant devices, augmented reality devices, virtual reality devices, and other devices that can send and receive data over network 102. Client devices 106 typically include user applications (such as a web browser) to facilitate sending and receiving data over network 102, but native applications (other than browsers) executed by client devices 106 may also facilitate sending and receiving data over network 102.

[0049] Game device is the device that enables the user to participate in game application, for example, in this device, the user can control one or more roles, avatar or other rendering content presented in the game application.Game device generally includes computer processor, memory device and the controller interface (physically or visually rendered) that enables the user to control the content rendered by the game application.Game device can store and execute game application locally, or execute the game application (for example, online game application) that is stored and / or served at least in part by cloud server. Similarly, game device can be with the game server docking that executes game application and " streams " game application to game device.Game device can be tablet device, mobile telecommunication device, computer or the another kind of device that also performs other functions except executing game application.

[0050] A digital assistant device includes a device having a microphone and a speaker. The digital assistant device is typically capable of receiving input via voice and responding to the content using audible feedback, and may present other audible information. In some cases, the digital assistant device also includes a visual display or communicates with a visual display (e.g., via a wireless or wired connection). When a visual display is present, feedback or other information may also be provided visually. In some cases, the digital assistant device may also control other devices, such as lights, locks, cameras, climate control devices, alarm systems, and other devices registered with the digital assistant device.

[0051] As shown, client device 106 is presenting electronic document 150. An electronic document is data that represents a set of content at client device 106. Examples of electronic documents include web pages, word processing documents, portable document format (PDF) documents, images, videos, search results pages, and feeds. Native applications (e.g., "apps" and / or game applications) (such as those installed on mobile, tablet, or desktop computing devices) are also examples of electronic documents. Electronic documents can be provided to client devices 106 by electronic document servers 104 ("Electronic Doc Servers").

[0052] For example, electronic document server 104 may comprise a server hosting a publisher website. In this example, client device 106 may initiate a request for a given publisher web page, and electronic server 104 hosting the given publisher web page may respond to the request by sending machine-executable instructions that initiate presentation of the given web page at client device 106.

[0053] In another example, the electronic document server 104 may include an app server from which the client device 106 can download the app. In this example, the client device 106 can download the files required to install the app at the client device 106 and then execute the downloaded app locally (e.g., on the client device). Alternatively or additionally, the client device 106 can initiate a request to execute the app, which is transmitted to the cloud server. In response to receiving the request, the cloud server can execute the application and stream the application's user interface to the client device 106, so that the client device 106 does not have to execute the app itself. Instead, the client device 106 can present the user interface generated by the cloud server executing the app and transmit any user interaction with the user interface back to the cloud server for processing.

[0054] An electronic document can include a variety of content. For example, an electronic document 150 may include native content 152 that resides within the electronic document 150 itself and / or does not change over time. An electronic document may also include dynamic content that may change over time or on a per-request basis. For example, the publisher of a given electronic document (e.g., electronic document 150) may maintain data sources used to populate various portions of the electronic document. In this example, the given electronic document may include a script, such as script 154, that causes the client device 106 (or cloud server) to request content (e.g., digital components) from the data source when the given electronic document is processed (e.g., rendered or executed) by the client device 106 (or cloud server). The client device 106 (or cloud server) integrates the content (e.g., digital components) obtained from the data source into the given electronic document to create a composite electronic document that includes the content obtained from the data source.

[0055] In some cases, a given electronic document (e.g., electronic document 150) may include a digital component script (e.g., script 154) that references service device 110 or a particular service provided by service device 110. In these cases, the digital component script is executed by client device 106 when the given electronic document is processed by client device 106. Execution of the digital component script configures client device 106 to generate a request for a digital component 112 (referred to as a "component request"), which is transmitted to service device 110 via network 102. For example, the digital component script may enable client device 106 to generate a packetized data request including header and payload data. Component request 112 may include event data specifying characteristics, such as the name (or network location) of the server for which the digital component is being requested, the name (or network location) of the requesting device (e.g., client device 106), and / or information that service device 110 may use to select one or more digital components or other content to provide in response to the request. Component request 112 is transmitted by client device 106 via network 102 (e.g., a telecommunications network) to a server at service device 110.

[0056] Component request 112 may include event data specifying other event characteristics, such as the characteristics of the electronic document being requested and the location of the electronic document at which the digital component may be presented. For example, event data specifying a reference (e.g., a URL) to an electronic document (e.g., a webpage) in which the digital component is to be presented, an available location of the electronic document for presenting the digital component, the size of the available location, and / or the type of media eligible for presentation at the location may be provided to service device 110. Similarly, event data specifying keywords associated with the electronic document ("document keywords" or simply "keywords") or entities referenced by the electronic document (e.g., people, places, or things) may also be included in component request 112 (e.g., as payload data) and provided to service device 110 to facilitate identifying digital components eligible for presentation with the electronic document. Event data may also include a search query submitted from client device 106 to obtain a search results page. For example, a query received by a user of the search system may be stored in query storage 111.

[0057] The component request 112 may also include event data related to other information, such as information that has been provided by the user of the client device, geographic information indicating the state or region where the component request is being submitted, or other information that provides the context of the environment in which the digital component will be displayed (e.g., the time of day of the component request, the day of the week of the component request, the type of device on which the digital component will be displayed (e.g., a mobile device or tablet device)). The component request 112 may be transmitted, for example, over a packetized network, and the component request 112 itself may be formatted as packetized data having a header and payload data. The header may specify the destination of the packet, and the payload data may include any of the information discussed above.

[0058] The service device 110 selects a digital component (e.g., third-party content such as video files, audio files, images, text, game content, augmented reality content, and combinations thereof, all of which may take the form of advertising content or non-advertising content) to be presented with a given electronic document (e.g., at a location specified by the script 154) in response to receiving the component request 112 and / or using information included in the component request 112.

[0059] In some implementations, the digital component is selected in less than a second to avoid errors that may result from delayed selection of the digital component. For example, a delay in providing the digital component in response to the component request 112 may cause page load errors at the client device 106 or cause portions of the electronic document to remain unpopulated even after other portions of the electronic document are presented at the client device 106.

[0060] Furthermore, as the delay in providing the digital component to the client device 106 increases, it becomes more likely that the electronic document will no longer be present at the client device 106 when the digital component is delivered to the client device 106, thereby negatively impacting the user's experience of the electronic document. Additionally, the delay in providing the digital component may cause the delivery of the digital component to fail, for example, if the electronic document is no longer present at the client device 106 when the digital component is provided.

[0061] In some implementations, the service device 110 is implemented in a distributed computing system that includes, for example, a server and a plurality of computing devices 114 that are interconnected and identify and distribute digital components in response to the request 112. The plurality of computing devices 114 operate together to select from millions of available digital components (DCs). 1-x) that are eligible for presentation in an electronic document. Millions of available digital components may be indexed, for example, in the digital component database 116. Each digital component index entry may reference a corresponding digital component and / or include distribution parameters (DP1 to DP2) that facilitate (e.g., trigger, constrain, or limit) the distribution / transmission of the corresponding digital component. x For example, the distribution parameters may facilitate (e.g., trigger) transmission of the digital component by requiring the component request to include at least one criterion that matches (e.g., completely or with some pre-specified level of similarity) one of the distribution parameters of the digital component.

[0062] In some implementations, the distribution parameters for a particular digital component may include a distribution keyword that must be matched (e.g., to an electronic document, a document keyword, or a term specified in the component request 112) in order for the digital component to be eligible for presentation. Additionally or additionally, the distribution parameters may include embeddings that may use various different data dimensions, such as website details and / or consumption details (e.g., page viewports, user scrolling speed, or other information about data consumption). The distribution parameters may also require that the component request 112 include information specifying a particular geographic region (e.g., a country or state) and / or information specifying that the component request 112 originated from a particular type of client device (e.g., a mobile device or tablet device) in order for the digital component to be eligible for presentation. The distribution parameters may also specify an eligibility value (e.g., a ranking score, or some other specified value) that is used to evaluate the eligibility of the digital component for distribution / transmission (e.g., with other available digital components).

[0063] Identifying eligible digital components can be divided into multiple tasks 117a through 117c, which are then assigned among computing devices within a group of multiple computing devices 114. For example, different computing devices in group 114 can each analyze different portions of digital component database 116 to identify various digital components having distribution parameters that match the information included in component request 112. In some implementations, each given computing device in group 114 can analyze different data dimensions (or groups of dimensions) and communicate (e.g., transmit) the results of the analysis (Res 1 through Res 3) 118a through 118c back to service device 110. For example, the results 118a through 118c provided by each of the computing devices in group 114 can identify a subset of digital components that are eligible for distribution in response to the component request and / or a subset of digital components having certain distribution parameters. Identifying the subset of digital components can include, for example, comparing event data with the distribution parameters and identifying a subset of digital components having distribution parameters that match at least some characteristics of the event data.

[0064] The service device 110 aggregates the results 118a to 118c received from a set of multiple computing devices 114 and uses information associated with the aggregated results to select one or more digital components to be provided in response to the request 112. For example, the service device 110 can select a set of winning digital components (one or more digital components) based on the results of one or more content evaluation processes, as described below. In turn, the service device 110 can generate and transmit response data 120 (e.g., digital data representing the response) over the network 102, which enables the client device 106 to integrate the set of winning digital components into a given electronic document so that the set of winning digital components (e.g., winning third-party content) and the content of the electronic document are presented together at a display of the client device 106.

[0065] In some implementations, the client device 106 executes instructions included in the reply data 120 that configure the client device 106 and enable it to obtain a set of winning digital components from one or more digital component servers 108. For example, the instructions in the reply data 120 may include a network location (e.g., a uniform resource locator (URL)) and a script that causes the client device 106 to transmit a server request (SR) 121 to the digital component server 108 to obtain a given winning digital component from the digital component server 108. In response to the request, the digital component server 108 will identify the given winning digital component specified in the server request 121 (e.g., within a database storing a plurality of digital components) and transmit digital component data (DC data) 122 to the client device 106, the DC data representing the given winning digital component in an electronic document at the client device 106.

[0066] When the client device 106 receives the digital component data 122, the client device will render the digital component (e.g., third-party content) and present the digital component at the location specified by or assigned to the script 154. For example, the script 154 may create a walled garden environment (such as a frame) that is presented within, e.g., next to, the native content 152 of the electronic document 150. In some implementations, the digital component is overlaid on (or adjacent to) a portion of the native content 152 of the electronic document 150, and the service device 110 may specify the presentation location within the electronic document 150 in the reply 120. For example, when the native content 152 includes video content, the service device 110 may specify a location or object within a scene depicted in the video content above which the digital component is to be presented.

[0067] The performance of keywords that led to the placement of digital components can be monitored. For example, data describing click-through rate, conversion rate, cost per click, and other data can be monitored for keywords associated with digital components. A particular keyword may have multiple performance data sets, where each set is associated with a particular digital component or group of digital components that was placed using that keyword. For example, a first digital component provider may have selected the keyword "hiking boots" to place a first set of digital components, while a second digital component provider may have selected the keyword "hiking boots" to place a second set of digital components. Therefore, the keyword "hiking boots" will have a first performance data set for the first set of digital components, and a second performance data set for the second set of digital components.

[0068] The service device 110 may also include an artificial intelligence system 160 configured to generate data for identifying digital components. As described in more detail herein, the artificial intelligence ("AI") system 160 may generate digital component selection data (keywords) for the digital components. Digital component providers may then use these keywords as selection criteria for their digital components. Keywords may be single words or terms, or a combination of a term and a number.

[0069] Large language models are models trained to generate and understand human language. LLMs are trained on large datasets of text and code, and they can be used for a variety of tasks. For example, LLMs can be trained to translate text from one language to another; summarize text, such as website content, search results, news articles, or research papers; answer questions about text, such as "What is the capital of Georgia?"; create chatbots that can converse with humans; and generate creative text, such as poetry, stories, and code.

[0070] Language model system 170 may include any suitable language model neural network that receives an input sequence of text tokens selected from a vocabulary and autoregressively generates an output sequence of text tokens from the vocabulary. For example, the language model in system 170 may be a Transformer-based language model neural network or a recurrent neural network-based language model.

[0071] In some cases, when the neural network used to implement the language model autoregressively generates an output sequence of tokens, the large language model in system 170 can be referred to as an autoregressive neural network. More specifically, the autoregressively generated output is created by generating each particular token in the output sequence conditioned on the current input sequence, including any tokens that precede the particular text token in the output sequence (i.e., tokens that have been generated for any previous positions in the output sequence that precede the particular position of the particular token), and contextual input that provides context for the output sequence.

[0072] For example, when generating a token at any given position in the output sequence, the current input sequence may include the token at any previous position in the input sequence and the output sequence that precedes the given position. As a specific example, the current input sequence may include the input sequence followed by the token at any previous position in the output sequence that precedes the given position. Alternatively, the input and current output sequences may be separated by one or more predetermined tokens within the current input sequence.

[0073] More specifically, to generate a specific word-gram at a specific position in the output sequence, the neural network of the language model system 170 may process the current input sequence to generate a score distribution (e.g., a probability distribution) that assigns a corresponding score (e.g., a corresponding probability) to each word-gram in the word-gram vocabulary. The neural network of the language model system 170 may then use the score distribution to select a word-gram from the vocabulary as the specific word-gram. For example, the neural network of the language model system 170 may greedily select the word-gram with the highest score, or may sample word-grams from the distribution, for example, using kernel sampling or another sampling technique.

[0074] As a specific example, the language model system 170 may include an autoregressive Transformer-based neural network that includes (i) multiple attention blocks, each of which applies a self-attention operation; and (ii) an output subnetwork that processes the output of the last attention block to generate a score distribution.

[0075] Large language models can have any of a variety of Transformer-based neural network architectures. Examples of such architectures include those described in: J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al., "Training compute-optimal large language models", arXiv preprint arXiv:2203.15556, 2022; J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, H. F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche, L. A. Hendricks, M. Rauh, P. Huang, A. Glaese, J. Welbl, S. Dathathri, S. Huang, J.Uesato, J. Mellor, I. Higgins, A. Creswell, N. McAleese, A.Wu, E. Elsen, SMJayakumar, E. Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M. Paganini, L.Sifre, L. Martens, XL Li, A. Kuncoro, A. Nematzadeh, E. Gribovskaya,D. Donato, A. Lazaridou, A. Mensch, J. Lespiau, M. Tsimpoukelli, N. Grigorev, D. Fritz, T. Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D. Toyama, C. de Masson d'Autume, Y. Li, T. Terzi, V. Mikulik, I. Babuschkin, A. Clark, D. de Las Casas, A. Guy, C.Jones,J. Bradbury, M. Johnson, B.A. Hechtman, L. Weidinger, I. Gabriel, W.S. Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving, "Scaling language models: Methods, analysis & insights from training gopher," CoRR, abs / 2112.11446, 2021; Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu, "Exploring the limits of transfer learning with a unified text-to-text transformer," (Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer), arXiv preprint arXiv:1910.10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, ApoorvKulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le, "Towards a human-like open-domain chatbot," CoRR, abs / 2001.09977, 2020; and Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell et al., “Language models are few-shot learners,” arXiv preprint arXiv:2005.14165, 2020. Examples of such models include multi-task unified models, zero-shot models, domain-specific models, or language representation models.

[0076] However, in general, a transformer-based neural network includes a sequence of attention blocks, and during processing a given input sequence, each attention block in the sequence receives a corresponding input hidden state for each input token in the given input sequence. The attention block then updates each of the hidden states at least in part by applying self-attention to generate a corresponding output hidden state for each of the input tokens. The input hidden state of the first attention block is the embedding of the input token in the input sequence, and the input hidden state of each subsequent attention block is the output hidden state generated by the previous attention block.

[0077] In this example, the output subnetwork processes the output hidden state generated by the last attention block in the sequence for the last input token in the input sequence to produce a score distribution.

[0078] Typically, because language models are autoregressive, service device 110 can use the same large language model in system 170 to generate multiple different candidate output sequences in response to the same request, for example, by using beam search decoding from the score distribution generated by the language model, using a sampling and sorting decoding strategy, using different random seeds for the pseudorandom number generator used when sampling different runs of the language model, or using another decoding strategy that exploits the autoregressive properties of the language model.

[0079] In some implementations, the language model in system 170 is pre-trained, i.e., trained on language modeling tasks that do not require providing evidence in response to user questions, and the service device 110 (e.g., using AI system 160) causes the language model in system 170 to generate an output sequence according to a predetermined grammar through natural language prompts in the input sequence.

[0080] Query and landing page specificity

[0081] In some implementations, the system 170 uses a large language model in a workflow that is used to generate a student model to predict performance for queries against a landing page. Once the student model is prepared and trained, the service 110 can provide the query and landing page as input 172 to the student model and receive a score or series of scores (e.g., a vector) that predicts performance for queries against the landing page as output 174.

[0082] As described above, incorporating a measure of query landing page specificity deviation can improve the overall performance of the student model.As will be described in later sections, the system 170 implements a complex workflow for generating specificity-aware student models in a manner that conserves physical computer resources and processing time.

[0083] Regarding query-to-landing page specificity, queries that are more specific than landing pages generally do not adversely affect the performance of query-to-landing page pairs. For example, such performance could be information need satisfaction, user engagement, or other measurable metrics that can be determined through traffic log analysis. However, as queries become more general and landing pages become more specific relative to each other, the actual performance of queries targeting landing pages tends to decline.

[0084] For example, consider a relatively specific landing page for fencing products, such as "Horse Fencing: Wooden and Electric," for the fairly general query "horse." While the query might score well on the landing page for a given relevance model, the landing page itself might not actually be a very good match due to the high generic-to-specific specificity deviation from the query to the landing page. This is because the page's topic might be too specific to the query and, therefore, might not satisfy the user's information need. Incorporating specificity awareness into the model can account for this relative difference in specificity.

[0085] Conversely, for the more specific query "non-slip rugs for kitchen and bathroom," for example, the landing page "Buy non-slip area rugs from ExtraStock" might be a good match, assuming the landing page has filtering options for kitchen and bathroom. However, a given relevance model might not account for the filtering options and, therefore, might predict that the query is less relevant to the landing page than it actually is. Once again, incorporating specificity awareness into the model accounts for this relative difference in specificity.

[0086] Workflow for building specificity-aware models

[0087] Figure 2 is a diagram of a workflow 200 for generating a specificity perception model. Figure 2 refer to Figures 3 to 6 Specifically, Figure 3 is a diagram of a workflow 300 for determining a document-to-document score for two documents, and Figure 4 is a diagram of a workflow 400 for determining a query-to-document score for a query and associated landing pages.

[0088] refer to Figure 5 A particular portion of the workflow 200 is depicted above the dashed line 201, which is a flow chart of an example process 500 for training a teacher model that determines a query-to-document score that is partially specificity-aware for an input of a query and a landing page selected for the query. Similarly, referring to Figure 6 Depicting a particular portion of the workflow 200 below the dashed line 201 is a flow diagram of an example process 600 for training a student model that determines a query-to-document score that is partially specificity-aware for an input of a query and a landing page selected for the query.

[0089] The workflow begins with a set of specifically labeled document-to-document pairs 202. The specifically labeled document-to-document pairs 202 constitute a training set for training a document-to-document large language model 210 that predicts the document-specific deviation from a first document (e.g., a landing page selected for a query) to a second document (e.g., a page referenced by a search result responsive to the query). The specifically labeled document-to-document pairs 202 can be generated, for example, by a human evaluator, and labels can be selected to quantify the specific deviation from the landing page to the document. For example, given a left / right pair, where left is the landing page and right is the document referenced by the search result:

[0090] [Landing_page] ↔ [Search_results_page]

[0091] The evaluator's task is to label the pair using one of the following:

[0092] Different (-)

[0093] The left side is more specific (-)

[0094] The right side is more specific (+)

[0095] Almost (+)

[0096] The first two labels have negative values, while the last two labels have positive values. Each document-to-document pair is labeled by N evaluators. From the labels, the system determines a numerical value that quantifies the deviation in specificity from the landing page document to the search result document.

[0097] Any suitable scoring method can be used. In one implementation, a majority voting method is used, where the score is given by the following formula:

[0098] If (num_pos ≥ num_neg), then D2D_score = 1, otherwise 0;

[0099] where num_pos is the number of positive values, and num_neg is the number of negative values.

[0100] In another implementation, cumulative differences are used, where the score is given by the following formula:

[0101] D2D_score = (num_pos - num_neg) / num_total

[0102] Where num_total is the total number of labels.

[0103] In some implementations, the score range can be normalized to be in the range of [0, 1].

[0104] Once the training set is generated, it is used to train the document-to-document large language model 210. The model 210 can be trained as described above. In some implementations, the document-to-document large language model 210 is trained to determine the probability of a set of N labels, where each label corresponds to a quantification of the document-to-document specific deviation. For example, Figure 3 As shown, a set 320 of N descriptive labels is selected for prediction by the model 210. Each label in the set 320 corresponds to a specificity deviation value V(c) within a given specificity deviation range. Once the document-to-document large language model 210 is trained, it determines a predicted value P(c) for each of the N labels for a given pair of documents 302 (e.g., documents D1 and D2). Each label prediction value P(c) is then multiplied by its corresponding label value, i.e., the specificity deviation value V(c), and these products are then added together to obtain an overall document-to-document score for the document pair D1 and D2. For example, assume that six labels, i.e., label 1...label 6, have corresponding label values 1.0, 0.8, 0.6, 0.4, 0.2, and 0.0, and for a given document pair D1 and D2, the model 210 generates corresponding label predictions 0.0, 0.1, 0.3, 0.2, 0.1, and 0.3. Therefore, the overall document-to-document score is:

[0105] 0.0*1.0 + 0.1*0.8 + 0.3*0.6 + 0.2*0.4 +0.2*0.2 + 0.3*0.0 = 0.38

[0106] The final document-to-document score is a quantification of the specificity of the first document (eg, a landing page for a query) relative to the specificity of the second document (eg, a search results page for search results responsive to the query). Other scoring formulas may also be used.

[0107] Once the document-to-document large language model 210 is trained, it is used to generate a training corpus of labeled query-to-document scores for query and document pairs. Each query-to-document score quantifies the query-to-landing page specificity deviation. Figure 4 As shown, a query-to-document score is determined for a given query 402 and landing page 404. To determine the query-to-document score for a query and landing page pair, a document-to-document score is first calculated. Thus, the document-to-document language model 210 receives data for the search result page linked to by each search result in a proper subset 410 of search results that are responsive to the query. The proper subset of search results are search results that are ranked higher than other search results in the set of search results determined to be responsive to the query. For example, in Figure 4 , the documents in proper subset 410 may correspond to the top five highest-ranked search results for query 402 .

[0108] For each of the search result pages in the proper subset 410, the document-to-document large language model 210 determines a document-to-document score for the landing page and the search result page. Once the document-to-document scores are determined, the system 170 determines a final query-to-document score based on the document-to-document scores. In some implementations, the query-to-document score is based on a central tendency value of the document-to-document scores. For example, assuming that Figure 4 The five document-to-document values determined for the landing page in [ 402 ] are 0.4, 0.2, 0.3, 0.8, and 0.9. If the central tendency is the average, the query-to-document score for query 402 and landing page 404 is 0.52. Other scoring formulas may also be used.

[0109] For a document pair of a landing page and a search result page, in some implementations, the input to the document-to-document large language model 210 is the title of the landing page and the title of the search result page. However, the document-to-document large language model 210 can also be trained on other or additional input data.

[0110] Now refer to Figure 5 and Figure 6The end-to-end process of the workflow 200 is described. The processes 500 and 600 may be performed by a data processing device and a memory storage system in data communication with each other, such as those described below with reference to Figure 9 The device described.

[0111] Process 500 begins with the system (e.g., system 170) selecting a query and landing page pair (502), and then the system performs the process described in sub-block 510. Sub-block 510 is repeated for each query and landing page pair to generate the data needed to construct the specifically labeled Q2D 222 for ensemble training of the teacher model 210.

[0112] For a query and landing page pair, process 500 obtains a proper subset of search results that are responsive to the query (512). A proper subset of search results is a search result that is ranked higher than other search results in the set of search results determined to be responsive to the query. Each search result is linked to a search result page. For example, process 500 may Figure 4 A query 402 is submitted to a search engine, and search results are received in response, and thus search result pages D_SR1 - D_SRM are accessed.

[0113] Then, the process 500 continues by pairing each of the search result pages with a landing page (514), and for each pairing, performing sub-block 516. In this way, the system 170 processes M pairs of landing pages and search result pages D_SR1-D_SRM, i.e., (landing page 404, D_SR1), ... (landing page 404, D_SRM).

[0114] For each search result page and landing page pair, the document-to-document large language model 210 generates tag probabilities, where each tag probability is generated for a specific tag in the tag set (518). Figure 3 As depicted in , the document-to-document large language model 210 generates a corresponding prediction value for each of the labels, ie, label1 . . . labelN.

[0115] Then, the process 500 determines a document-to-document score for the landing page and search result page pair based on the label probabilities, which is a quantification of the specificity of the landing page relative to the specificity of the search result page, i.e., the specificity deviation from the landing page to the search result page (520). The document-to-document score can be as described above with reference to Figure 3 Described generation.

[0116] After determining the document-to-document score for each of the search result and landing page pairs for a given landing page, the process then determines a query-to-document score based on the document-to-document score that quantifies the query-to-landing page specificity deviation (522). For example, Figure 4 As shown, for the M document-to-document scores of the search result page 410, a query-to-document score is determined for the query 402 and the landing page 404. The query-to-document score can be as described above with reference to Figure 4 Described generation.

[0117] The process 500 then stores the training tuples in a teacher model-specific training set, where each training tuple is a query and landing page pair and a query-to-document score (524). Figure 2 In

[0045] , the teacher model-specific training set is specifically labeled query-document pairs 212. In some implementations, tens of millions of query-landing page pairs can be labeled for the training set 212.

[0118] The process 200 then trains a query-to-document teacher model 220 on the teacher model training set to determine a query-to-document score for the query input and the landing page selected for the query (526). Once trained, the teacher model 220 determines the query-to-document score without the search results and search result pages determined to be responsive to the query. More specifically, the teacher model 200 receives data from the query and the landing page to determine the query-to-landing page score. Thus, by using the teacher model 220, requests for search results and multiple inferences of the document-to-document large language model 210 are eliminated, thereby saving significant physical resources and reducing processing time.

[0119] The teacher model 220 is then used to generate a student model-specific training set 222, which is then used for ensemble training of the student model 230. Figure 6 This is described in the figure which depicts an example process 600 for training a student model that determines a query-to-document score that is, in part, specificity-aware for an input of a query and a landing page selected for the query.

[0120] Process 600 uses the teacher model 220, after the teacher model has been trained, to process query and landing page pairs to determine a query-to-document score for each query and landing page pair (602). The teacher model 220 receives only the query and landing page pair as input and does not need to obtain search results responsive to the query or determine a separate landing page for the search result document score. Therefore, less computer resources and time are required to determine a query-to-document-specific score for a query and landing page pair.

[0121] The process 600 stores training tuples in a student model-specific training set, each training tuple being a query and landing page pair and a query-to-document score determined for the query and landing page pair (604). Figure 2 , the student model-specific training set is specifically labeled query-document pairs 222. In some implementations, the training set 222 can include up to one billion or more labeled training pairs.

[0122] The process 600 uses the teacher model training set and at least one other student model training set that does not quantify the query-to-landing page specific deviation to perform ensemble training on the student model to determine a query-to-document score for the query input and the landing page selected for the query (606). The ensemble training can be completed using other labeled query-document data 224, such as data labeled with click quality (CQ) for query-to-landing page pairs, data labeled with information quality (IQ) for query-to-landing page pairs, and data labeled with landing page quality (LQ) for query-to-landing page pairs. The click quality labeled data quantifies the click probability of the document provided in response to the query. The information quality quantifies the information satisfaction of the document with respect to the query. The landing page quality quantifies the quality of the landing page's metric with respect to the query. Other labeled data derived from traffic logs can also be used for ensemble training.

[0123] In some implementations, ensemble training is performed such that when other labeled data is uncertain about the quality of the landing page and the performance of the query pair, the specifically labeled data is weighted more heavily. For example, for a given landing page and query pair, when both the information quality (IQ) labeled data and the landing page quality (LQ) labeled data indicate scores that exceed corresponding high score thresholds (i.e., each corresponding score exceeds a corresponding "high" score threshold), the training loss of the sample is increased. Conversely, for a given landing page and query pair, when both the information quality (IQ) labeled data and the landing page quality (LQ) labeled data indicate scores that do not meet corresponding low score thresholds (e.g., each corresponding score is less than a corresponding "low" score threshold), the training loss of the sample is reduced.

[0124] However, if neither of the above two conditions is met, the specificity-labeled data will be weighted as being most indicative of performance. Therefore, when the combination of information quality (IQ) and landing page quality (LQ) is not indicative of good or bad performance, that is, when the combination of information quality (IQ) labels and landing page quality (LQ) labels may be uncertain for performance prediction, ensemble training produces a specificity-aware and more specificity-reliant student model.

[0125] Once the student model 230 is trained, it can be used by the service device 110 as described above. Figure 7Additional details regarding the teacher model 220 are described and referenced below. Figure 8 Additional details regarding the student model 230 are described.

[0126] Model Architecture

[0127] As described above, the document-to-document large language model 220 using the decoder requires a separate inference for each label prediction when determining the document-to-document score. Therefore, since each landing page has M documents scored based on specificity, and each document requires N labels to be predicted, the document-to-document large language model requires N×M inferences to determine the query-to-document score for each landing page.

[0128] In another implementation, the document-to-document large language model 220 is an encoder-only model. An example encoder-only architecture is the EncT5 encoder model. When using an encoder-only model, N separate inferences are not required to generate the corresponding N predictions for each of the labels 1-N. Instead, the label probabilities for the N labels are determined by a single call. Therefore, the processing time for generating the training set 212 can be significantly reduced. In addition, since the decoding process step is removed, the processing time for generating the predictions is further reduced. In operation, the output of the encoder includes a logistic regression (logit) for each of the label predictions, and the logistic regression is processed directly to generate the entire set of label predictions in a single pass.

[0129] Figure 7 is an example architecture of the teacher model 220. For example, the teacher model 220, which can be a cross-attention model, is trained on one or more query dimensions and one or more document dimensions, such as the query field 710 and the document field 720. For example, the field 710 includes text, salient terms, geographic location, and entities. The text field includes the text of the query. Salient terms, which are terms derived from the query, are terms that are determined to be of topical importance. The geographic location is the location specified in the query and / or the location from which the query is received. For the document field 720, the URL field is the uniform resource locator of the document, and the title field includes the title text of the document. Salient terms are terms that are derived from the document and are determined to be of topical importance.

[0130] The entity fields of the query field 710 and the document field 720 are knowledge graph entities determined to be relevant to the query. A knowledge graph is a graph of entities, where each entity is represented by a node and edges between entities indicate that the entities are related. A knowledge graph typically represents a network of real-world entities (such as objects, events, or concepts) and the relationships between entities.

[0131] For queries and documents, the service device 110 can determine entities based on an analysis of the query text and document text, locations associated with the query and document, and other data related to the query and document. For the query and document, the system 100 determines the average number of entities determined for each and considers the percentile ranking of the number of entities determined. The system then selects the number of entity features for each of the query field and the document field. In some implementations, the query entity features are selected as 8 and the document entity features are selected as 16. These values provide an acceptable trade-off between training and inference time degradation and performance improvement.

[0132] The input is tokenized by the tokenizer into a sequence of tokens 730, such as an input integer. The embedding layer then takes the tokens and generates a combined embedding 740 by mapping the tokens to embeddings. Tokenization is the process of arranging the input data into tokens that can be mapped into a vector space.

[0133] Transformer layer 750 receives the combined embeddings as input. The combined embeddings are derived from two or more query dimensions and two or more document embeddings for the input query and landing page pair. Transformer layer 750 maps the input sequence of embeddings into a sequence of continuous representations. The output of transformer layer 750 is provided to pooler layer 760, which in turn produces a vector representation of the input sequence. The generated vector represents the query-to-document score data for the input query and landing page pair.

[0134] Figure 8 is an example architecture of the student model 230. In some implementations, the student model 230 is a dual encoder and includes a query encoder 810 and a document encoder 850. The query encoder 810 generates a query embedding from the combined embeddings of two or more query dimensions, and the document encoder 850 generates a document embedding from the combined embeddings of two or more document embeddings.

[0135] The student model 230 is integratedly trained on one or more query dimensions and one or more document dimensions, such as the query field 812 and the document field 852. For example, field 812 includes text, significant terms, geographic location, and entities. The text field includes the text of the query. Significant terms are terms derived from the query and determined to be of topical importance. The geographic location is the location specified in the query and / or the location from which the query was received. For the document field 852, the URL field is the uniform resource locator of the document, and the title field includes the title text of the document. Significant terms are terms derived from the document and determined to be of topical importance.

[0136] The input is a sequence of tokens 814 and 854, for example, an input integer, which are tokenized by the tokenizer. Then, the corresponding embedding layer takes the corresponding tokens 814 and 854 and generates the corresponding combined embeddings 816 and 856 by mapping the corresponding tokens 814 and 854 to the corresponding embeddings 816 and 856.

[0137] The respective transformer layers 818 and 858 receive the respective combined embeddings 818 and 858 as input and map the embedded input sequence to a series of continuous representations. The outputs of the respective transformer layers 818 and 858 are provided to the respective pooler layers 820 and 860, which are query embeddings 822 and document embeddings 862, respectively. The query embeddings 822 and document embeddings 862 are used to predict the performance of the input query of the input landing page. For example, a cosine similarity score that measures the similarity between the query embedding 822 and the document embedding 862 can be used to predict the score. Other suitable scoring processes can also be used.

[0138] Figure 7 and Figure 8 The model can be trained on any set of query and document fields shown. For example, the model can be trained on any combination of text, significant terms, geolocation, URL, and entity fields. For example, the model can only be trained on text, significant terms, geolocation, and URL fields. In addition, Figure 7 and Figure 8 The other fields shown in are now available for training as well.

[0139] Additional implementation details

[0140] Figure 9 9 is a block diagram of an example computer system 900 that can be used to perform the operations described above. System 900 includes a processor 910, a memory 920, a storage device 930, and an input / output device 940. Each of components 910, 920, 930, and 940 can be interconnected, for example, using a system bus 950. Processor 910 is capable of processing instructions for execution within system 900. In one implementation, processor 910 is a single-threaded processor. In another implementation, processor 910 is a multi-threaded processor. Processor 910 is capable of processing instructions stored in memory 920 or on storage device 930.

[0141] The memory 920 stores information within the system 900. In one implementation, the memory 920 is a computer-readable medium. In one implementation, the memory 920 is a volatile memory unit. In another implementation, the memory 920 is a non-volatile memory unit.

[0142] The storage device 930 can provide mass storage for the system 900. In one implementation, the storage device 930 is a computer-readable medium. In various implementations, the storage device 930 may include, for example, a hard disk device, an optical disk device, a storage device shared by multiple computing devices over a network (e.g., a cloud storage device), or some other mass storage device.

[0143] The input / output device 940 provides input / output operations for the system 900. In one implementation, the input / output device 940 may include one or more of a network interface device (e.g., an Ethernet card), a serial communication device (e.g., an RS-232 port), and / or a wireless interface device (e.g., an 802.11 card). In another implementation, the input / output device may include a driver device configured to receive input data and send output data to other devices (e.g., a keyboard, a printer, a display, and other peripheral devices 960). However, other implementations may also be used, such as mobile computing devices, mobile communication devices, set-top television client devices, etc.

[0144] Despite Figure 9 An example processing system is described in the specification, but the subject matter and implementation of the functional operations described in this specification may be implemented in other types of digital electronic circuit systems or in computer software, firmware or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them.

[0145] An electronic document (referred to simply as a document for brevity) does not necessarily correspond to a file. A document may be stored as part of a file that holds other documents, in a single file dedicated to the document in question, or in multiple coordinated files.

[0146] With respect to situations in which the systems discussed herein collect and / or use personal information about users, users may be provided with the opportunity to enable / disable or otherwise control programs or features that may collect and / or use personal information (e.g., information about the user's social network, social actions or activities, the user's preferences, or the user's current location). Additionally, certain data may be processed in one or more ways prior to its storage or use so that personally identifiable information associated with the user is removed. For example, the user's identity may be anonymized so that personally identifiable information about the user cannot be determined, or, where location information is available, the user's geographic location may be generalized (such as to a city, zip code, or state level) so that the user's specific location cannot be determined.

[0147] The embodiments of the subject matter and operations described in this specification may be implemented in digital electronic circuit systems or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them. The embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, which are encoded on a computer storage medium for execution by a data processing device or for controlling the operation of a data processing device. Alternatively or in addition, the program instructions may be encoded on an artificially generated propagation signal (e.g., a machine-generated electrical, optical, or electromagnetic signal), which is generated to encode information for transmission to a suitable receiver device for execution by the data processing device. The computer storage medium may be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them, or included therein. Furthermore, although a computer storage medium is not a propagation signal, a computer storage medium may be the source or destination of computer program instructions encoded in an artificially generated propagation signal. The computer storage medium can also be, or be included in, one or more separate physical components or media (eg, multiple CDs, disks, or other storage devices).

[0148] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0149] The term "data processing equipment" includes all types of equipment, devices and machines for processing data, for example, including programmable processors, computers, systems on a chip, or a plurality of or a combination thereof. The equipment may include a dedicated logic circuit system, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, the equipment may also include code that creates an execution environment for the computer program in question, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of these. The equipment and execution environment can implement various different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.

[0150] This document relates to service devices. As used herein, a service device is one or more data processing devices that perform operations to facilitate the distribution of content over a network. A service device is depicted as a single block in the block diagram. However, while a service device can be a single device or a single group of devices, the present disclosure contemplates that a service device can also be a group of devices, or even multiple different systems that communicate to provide various content to client devices. For example, a service device can include one or more of a search system, a video streaming service, an audio streaming service, an email service, a navigation service, an advertising service, a gaming service, or any other service.

[0151] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages or declarative or procedural languages; and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network.

[0152] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by, and devices can be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0153] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or be operatively coupled to receive data from or transfer data to, or both. However, a computer need not have such devices. Furthermore, a computer may be embedded in another device, such as a mobile phone, personal digital assistant (PDA), mobile audio or video player, game console, global positioning system (GPS) receiver, or portable storage device (e.g., universal serial bus (USB) flash drive), to name a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0154] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound, voice, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on a user's client device in response to a request received from the web browser.

[0155] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, such as a data server, or includes a middleware component, such as an application server, or includes a front-end component, such as a client computer having a graphical user interface or a web browser through which a user can interact with implementations of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).

[0156] A computing system may include a client and a server. The client and server are typically remote from each other and typically interact via a communication network. The relationship of client and server arises from computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server transmits data (e.g., an HTML web page) to a client device (e.g., for the purpose of displaying data to a user interacting with the client device and receiving user input from the user). Data generated at the client device (e.g., the results of the user interaction) may be received from the client device at the server.

[0157] Although this specification contains many specific implementation details, these details should not be interpreted as limiting the scope of any invention or the scope that may be claimed, but should be interpreted as a description of features that are unique to a particular embodiment of a particular invention. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. In addition, although features may be described above as working in certain combinations and even initially claimed as such, in some cases, one or more features from the claimed combination may be deleted from the combination, and the claimed combination may involve a sub-combination or a variation of the sub-combination.

[0158] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in a continuous order, or that all illustrated operations be performed, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0159] Thus, certain embodiments of the present subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. Additionally, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing may be advantageous.

Claims

1. A computer-implemented method comprising: For each of the plurality of query and landing page pairs, each query and landing page pair is a specific query and a specific landing page selected for the query: For the query, obtaining a proper subset of search results responsive to the query, the proper subset of search results being search results determined to be ranked higher than other search results in the set of search results responsive to the query, wherein each search result is linked to a search results page; For each search results page and said landing page: generating a plurality of label probabilities by a large language model trained to determine label probabilities, each label probability being generated for a particular label in the label set and generated by a separate inference call of the large language model; determining a document-to-document score from the plurality of label probabilities, the document-to-document score being a quantification of the specificity of the landing page relative to the specificity of the search result page; determining a query-to-document score from the document-to-document scores, the query-to-document score quantifying a query-to-landing page specificity deviation; as well as storing training tuples in a teacher model-specific training set, each training tuple being a query and landing page pair and the query-to-document score determined for the query and landing page pair; as well as A teacher model is trained on the teacher model training set to determine a query-to-document score for an input of a query and a landing page selected for the query, wherein the query-to-document score is determined without search results and search result pages determined to be responsive to the query.

2. The method of claim 1, further comprising: processing a plurality of query and landing page pairs using the teacher model after the teacher model is trained to determine a query-to-document score for each query and landing page pair; as well as Training tuples are stored as a student model-specific training set, each training tuple being a query and landing page pair and the query-to-document score determined for the query and landing page pair.

3. The method of claim 2, further comprising: The student model is ensemble trained to predict performance for queries targeting landing pages using the student model-specific training set and at least one other student model training set that does not quantify the query-to-landing page specific deviation.

4. The method of claim 3, wherein performing ensemble training on the student model comprises performing ensemble training on a dual encoder.

5. The method of claim 4, wherein the dual encoder comprises: a query encoder that generates a query embedding from the combined embeddings of two or more query dimensions; as well as A document encoder that generates a document embedding from the combined embeddings of two or more document dimensions.

6. The method of claim 3, wherein the teacher model comprises a crisscross attention model, the crisscross attention model comprising: a transformer layer that receives as input the combined embeddings from the two or more query dimensions and two or more document embeddings for an input query and landing page pair; as well as A pooler layer receives the output of the transformer layer and generates query-to-document score data for the input query and landing page pair.

7. The method of claim 3, wherein the other student model training set includes data labeled by information quality (IQ) and landing page quality (LQ), and the ensemble training comprises: When the data marked by the information quality IQ and the data marked by the landing page quality LQ both indicate scores exceeding corresponding high score thresholds, increasing the training loss of the sample; as well as When both the data marked by the information quality IQ and the data marked by the landing page quality LQ indicate scores that do not meet corresponding low score thresholds, the training loss of the sample is reduced.

8. The method of claim 1 , wherein determining the query-to-document score that quantifies the query-to-landing page specificity deviation from the document-to-document score comprises: A central tendency value is determined for the document-to-document scores determined for the query and landing page pair.

9. A system comprising: one or more computers in data communication; as well as One or more non-transitory computer-readable media storing instructions executable by the one or more computers, the instructions, when executed by the one or more computers, causing the one or more computers to perform operations comprising: For each of the plurality of query and landing page pairs, each query and landing page pair is a specific query and a specific landing page selected for the query: For the query, obtaining a proper subset of search results responsive to the query, the proper subset of search results being search results determined to be ranked higher than other search results in the set of search results responsive to the query, wherein each search result is linked to a search results page; For each search results page and said landing page: generating a plurality of label probabilities by a large language model trained to determine label probabilities, each label probability being generated for a particular label in the label set and generated by a separate inference call of the large language model; determining a document-to-document score from the plurality of label probabilities, the document-to-document score being a quantification of the specificity of the landing page relative to the specificity of the search result page; determining a query-to-document score from the document-to-document scores, the query-to-document score quantifying a query-to-landing page specific deviation; and storing training tuples in a teacher model-specific training set, each training tuple being a query and landing page pair and the query-to-document score determined for the query and landing page pair; and A teacher model is trained on the teacher model training set to determine a query-to-document score for an input of a query and a landing page selected for the query, wherein the query-to-document score is determined without search results and search result pages determined to be responsive to the query.

10. The system of claim 9, the operations further comprising: processing a plurality of query and landing page pairs using the teacher model after the teacher model is trained to determine a query-to-document score for each query and landing page pair; as well as Training tuples are stored as a student model-specific training set, each training tuple being a query and landing page pair and the query-to-document score determined for the query and landing page pair.

11. The system of claim 10, wherein the operations further comprise: The student model is ensemble trained to predict performance for queries targeting landing pages using the student model-specific training set and at least one other student model training set that does not quantify the query-to-landing page specific deviation.

12. The system of claim 11, wherein ensemble training the student model comprises ensemble training a dual encoder.

13. The system of claim 12, wherein the dual encoder comprises: a query encoder that generates a query embedding from the combined embeddings of two or more query dimensions; as well as A document encoder that generates a document embedding from the combined embeddings of two or more document dimensions.

14. The system of claim 11, wherein the teacher model comprises a crisscross attention model, the crisscross attention model comprising: a transformer layer that receives as input the combined embeddings from the two or more query dimensions and two or more document embeddings for an input query and landing page pair; as well as A pooler layer receives the output of the transformer layer and generates query-to-document score data for the input query and landing page pair.

15. The system of claim 11, wherein the other student model training set includes data labeled with information quality (IQ) and landing page quality (LQ), and the ensemble training comprises: When the data marked by the information quality IQ and the data marked by the landing page quality LQ both indicate scores exceeding corresponding high score thresholds, increasing the training loss of the sample; as well as When both the data marked by the information quality IQ and the data marked by the landing page quality LQ indicate scores that do not meet corresponding low score thresholds, the training loss of the sample is reduced.

16. The system of claim 9, wherein determining the query-to-document score that quantifies the query-to-landing page specificity deviation from the document-to-document score comprises: A central tendency value is determined for the document-to-document scores determined for the query and landing page pair.

17. One or more non-transitory computer-readable media storing instructions that, when executed by an artificial intelligence system, cause the artificial intelligence system to perform operations comprising: For each of the plurality of query and landing page pairs, each query and landing page pair is a specific query and a specific landing page selected for the query: For the query, obtaining a proper subset of search results responsive to the query, the proper subset of search results being search results determined to be ranked higher than other search results in the set of search results responsive to the query, wherein each search result is linked to a search results page; For each search results page and said landing page: generating a plurality of label probabilities by a large language model trained to determine label probabilities, each label probability being generated for a particular label in the label set and generated by a separate inference call of the large language model; determining a document-to-document score from the plurality of label probabilities, the document-to-document score being a quantification of the specificity of the landing page relative to the specificity of the search result page; determining a query-to-document score from the document-to-document scores, the query-to-document score quantifying a query-to-landing page specificity deviation; as well as storing training tuples in a teacher model-specific training set, each training tuple being a query and landing page pair and the query-to-document score determined for the query and landing page pair; as well as A teacher model is trained on the teacher model training set to determine a query-to-document score for an input of a query and a landing page selected for the query, wherein the query-to-document score is determined without search results and search result pages determined to be responsive to the query.

18. The non-transitory computer-readable medium of claim 17, the operations further comprising: processing a plurality of query and landing page pairs using the teacher model after the teacher model is trained to determine a query-to-document score for each query and landing page pair; as well as Training tuples are stored as a student model-specific training set, each training tuple being a query and landing page pair and the query-to-document score determined for the query and landing page pair.

19. The non-transitory computer-readable medium of claim 18, the operations further comprising: The student model is ensemble trained to predict performance for queries targeting landing pages using the student model-specific training set and at least one other student model training set that does not quantify the query-to-landing page specific deviation.