Selecting digital components using intermediate embedding of language model neural network

By selecting and presenting digital components using intermediate embeddings generated by language model neural networks, the problems of delay and error in the prior art are solved, and more flexible and accurate selection and presentation of digital components are achieved.

CN120019378APending Publication Date: 2025-05-16GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072015.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-07-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art has the possibility of delay and error when selecting and presenting digital components related to machine learning model output, and it is difficult to flexibly and accurately analyze the context of the prompt to select the appropriate digital components.

Method used

The target digital components are selected by using the language model neural network to determine the target phase and select the corresponding digital components by using the intermediate embedding generated when processing natural language prompts, and the distance between the intermediate embedding and the reference embedding is calculated in the latent space to determine the target phase and select the corresponding digital components.

Benefits of technology

Reduces the delay in selecting additional content, reduces the possibility of errors, and improves flexible and accurate analysis of natural language prompt contexts, thereby more efficiently selecting and presenting digital components.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019378A_ABST
    Figure CN120019378A_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for selecting digital components using intermediate embedding of a language model neural network are described. A natural language cue is received from a user, and the natural language cue is processed by a language model neural network to generate a response. And obtaining intermediate embedding. One or more target phases are determined from among a plurality of different phases included in a user's journey during which the user performs different actions to perform a computer-implemented task based on the intermediate embedding. One or more target digital components are selected from the respective candidate digital components mapped to the one or more target stages. One or more target digital components are presented for display to the user along with a response to the natural language cue that has been generated by the language model neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] This specification relates to machine learning.

[0002] A machine learning model receives input and generates an output, such as a predicted output, based on the received input. Some machine learning models are parametric models and generate output based on the received input and the values ​​of the model parameters.

[0003] Some machine learning models are deep models that use multiple layers of models to generate outputs for received inputs. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a nonlinear transformation to a received input to generate an output. Summary of the invention

[0004] The present disclosure describes selecting digital components using intermediate embeddings of a language model neural network. A system can be implemented as a computer program on one or more computers at one or more locations that uses intermediate embeddings of a language model neural network generated during processing a prompt to generate an output to select a target digital component for presentation. For example, the language model neural network can be a Transformer-based language model neural network or a recursive neural network-based language model.

[0005] According to one aspect, a method performed by one or more computers is provided, the method comprising: receiving a natural language prompt from a user; processing the natural language prompt by a language model neural network to generate a response to the natural language prompt; obtaining an intermediate embedding generated by the language model neural network during processing the natural language prompt to generate the response; determining one or more target stages from a plurality of different stages included in a user journey based on the intermediate embedding, during which the user performs different actions to perform a computer-implemented task; selecting one or more target digital components from corresponding candidate digital components mapped to the one or more target stages; and presenting the one or more target digital components and the response to the natural language prompt that has been generated by the language model neural network for display to the user.

[0006] Other embodiments of this aspect include corresponding systems and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the above method aspects.

[0007] The foregoing and other embodiments may each optionally include one or more of the following features, alone or in combination. In particular, one embodiment includes all of the following features in combination.

[0008] The computer-implemented task may include navigating from a source web page through a plurality of web pages to reach a landing web page, wherein the landing web page represents a solution to a problem represented by the source web page.

[0009] There may be a progression of different stages including a problem awareness stage, followed by a solution provider awareness stage, followed by a solution consideration stage, followed by a solution comparison stage, followed by a solution implementation stage.

[0010] The intermediate embeddings may include: the output of an intermediate neural network layer of a language model neural network.

[0011] Determining one or more target stages from among a plurality of different stages included in a user journey may include: calculating a distance between an intermediate embedding and each of a plurality of reference embeddings in a latent space, wherein the plurality of reference embeddings are generated by a language model neural network based on processing historical natural language prompts provided by different users at different reference stages included in a corresponding reference user journey; selecting one or more reference embeddings having a distance that satisfies a distance threshold; identifying one or more reference stages within the corresponding reference user journeys during which the one or more reference embeddings were generated by the language model neural network; and using the one or more reference stages as one or more target stages.

[0012] The method may include: using a language model neural network to generate additional historical natural language prompts based on information available on the landing web page, the additional historical natural language prompts representing historical natural language prompts to be provided by different users at different reference stages included in corresponding reference user journeys.

[0013] Selecting one or more target digital components may include receiving, from a digital component provider and for each of a plurality of different stages included in the user journey, a respective candidate digital component associated with each stage.

[0014] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.

[0015] By using an intermediate embedding of a language model neural network that has been generated during processing of a prompt to generate an output, selection of additional content (e.g., a digital component) to present with the output can be made much faster than selecting the additional content based on subsequent analysis of the output after the output is generated, i.e., because the selection of the additional content can be performed at least in part concurrently with the generation of the output. This reduces latency in providing and presenting the content, which can also reduce errors that may occur while waiting for the additional content.

[0016] By measuring the distance between intermediate embeddings and reference embeddings that are mapped to known stages within the corresponding user journey in the latent space, the system can more flexibly and accurately analyze the context of the prompt, which facilitates the identification of more appropriate digital components for presentation. In some examples, these digital components, when presented to the user, will help the user, for example, in subsequent conversation turns to provide prompts in a manner that can reduce the total number of conversation turns that would otherwise be required to complete a particular task. Therefore, the costs associated with the computational resources (e.g., memory, storage, data transfer, and computing power) required for the language model to process additional prompts are also reduced. Further, the likelihood of displaying inappropriate content in response to a prompt can be reduced without introducing significant delays in displaying appropriate content.

[0017] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a block diagram of an example environment in which one or more target digital components may be presented according to implementations of the present disclosure.

[0019] Figure 2 is an example illustration of a user journey for performing a computer-implemented task according to implementations of the present disclosure.

[0020] Figure 3 Example operations for selecting a target digital component for presentation according to implementations of the present disclosure are shown.

[0021] Figure 4 is a flow chart of an example process for selecting a target digital component and presenting the target digital component along with a response to a natural language prompt according to an implementation of the present disclosure.

[0022] Figure 5 According to the implementation method of the present disclosure Figure 4 A flowchart of a process step with sub-steps within a step.

[0023] Figure 6 is a block diagram of an example computer system that may be used to perform the operations described herein in accordance with implementations of the present disclosure.

[0024] Like reference numbers and designations in the various drawings indicate like elements. DETAILED DESCRIPTION

[0025] The following detailed description describes the use of intermediate embeddings of language model neural networks to select digital components, and is presented to enable any technician in the field to make and use the disclosed subject matter in the context of one or more specific implementations. Various modifications, changes, and substitutions may be made to the disclosed implementations without departing from the scope of the present disclosure, and these modifications, changes, and substitutions are obvious to those of ordinary skill in the art, and the general principles defined may be applied to other implementations and applications. In some cases, one or more technical details that are not necessary to obtain an understanding of the described subject matter and are within the technical scope of those of ordinary skill in the art may be omitted without obscuring one or more described implementations. The present disclosure is not intended to be limited to the implementations described or shown, but rather to the widest scope consistent with the described principles and features.

[0026] Figure 1 is a block diagram of an example environment in which one or more target digital components may be presented according to an implementation of the present disclosure. As used throughout this document, the phrase "digital component" refers to a discrete unit of digital content or digital information (e.g., a video clip, an audio clip, a multimedia clip, game content, an image, text, a bullet point, an artificial intelligence output, a language model output, or another content unit). A digital component may be electronically stored in a physical memory device as a single file or as a collection of files, and a digital component may take the form of a video file, an audio file, a multimedia file, an image file, or a text file and include advertising information, such that an advertisement is a type of digital component.

[0027] The example environment 100 includes a network 102, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof. The network 102 connects electronic document servers 104, client devices 106, digital component servers 108, and service devices 110. The example environment 100 may include many different electronic document servers 104, client devices 106, and digital component servers 108.

[0028] Client devices 106 are electronic devices capable of requesting and receiving online resources over network 102. Example client devices 106 include personal computers, gaming devices, mobile communication devices, digital assistant devices, wearable devices (e.g., smart watches), augmented reality devices, virtual reality devices, and other devices that can send and receive data over network 102. Client devices 106 typically include user applications (such as web browsers) to facilitate sending and receiving data over network 102, but native applications (other than browsers) executed by client devices 106 may also facilitate sending and receiving data over network 102.

[0029] A game device is a device that enables a user to participate in a game application, for example, in which a user can control one or more characters, avatars or other rendered content presented in a game application. A game device generally includes a computer processor, a memory device and a controller interface (physical or visually rendered) that enables a user to control the content rendered by the game application. The game device can store and execute the game application locally, or execute a game application (for example, an online game application) stored and / or served at least in part by a cloud server. Similarly, the game device can be docked with a game server that executes the game application and "streams" the game application to the game device. The game device can be a tablet device, a mobile telecommunications device, a computer or another device that performs other functions in addition to executing the game application.

[0030] The digital assistant device includes a device with a microphone and a speaker. The digital assistant device is generally capable of receiving input by voice, and responding to the content using audible feedback, and can present other audible information. In some cases, the digital assistant device also includes a visual display or communicates with a visual display (e.g., by wireless or wired connection). When there is a visual display, feedback or other information can also be provided visually. In some cases, the digital assistant device can also control other devices, such as lights, locks, cameras, climate control devices, alarm systems, and other devices registered with the digital assistant device.

[0031] As shown, client device 106 is presenting electronic document 150. An electronic document is data that presents a set of content at client device 106. Examples of electronic documents include the output of a language model, a web page, a word processing document, a portable document format (PDF) document, an image, a video, a search result page, and a feed. Native applications (e.g., "apps" and / or game applications) such as applications installed on a mobile, tablet, or desktop computing device are also examples of electronic documents. Electronic documents can be provided to client device 106 by electronic document server 104 ("E-Doc Server").

[0032] For example, electronic document server 104 may include a server hosting a publisher website. In this example, client device 106 may initiate a request for a given publisher web page, and electronic server 104 hosting a given publisher web page may respond to the request by sending a machine executable instruction that initiates presentation of the given web page at client device 106.

[0033] In another example, the electronic document server 104 may include an app server from which the client device 106 may download the app. In this example, the client device 106 may download the files required to install the app at the client device 106, and then execute the downloaded app locally (i.e., on the client device). Alternatively or in addition, the client device 106 may initiate a request to execute the application, which is transmitted to the cloud server. In response to receiving the request, the cloud server may execute the application and stream the user interface of the application to the client device 106, so that the client device 106 does not have to execute the app itself. Instead, the client device 106 may present the user interface generated by the cloud server executing the app, and transmit any user interaction with the user interface back to the cloud server for processing.

[0034] An electronic document may include a variety of content. For example, an electronic document 150 may include native content 152 that is located within the electronic document 150 itself and / or does not change over time. An electronic document may also include dynamic content that may change over time or upon request. For example, a publisher of a given electronic document (e.g., electronic document 150) may maintain data sources for populating various parts of the electronic document. In this example, a given electronic document may include a script, such as script 154, that causes the client device 106 to request content (e.g., digital components) from a data source when the given electronic document is processed (e.g., rendered or executed) by a client device 106 (or a cloud server). The client device 106 (or cloud server) integrates the content (e.g., digital components) obtained from the data source into a given electronic document to create a composite electronic document including the content obtained from the data source.

[0035] In some cases, a given electronic document (e.g., electronic document 150) includes a script (e.g., script 154) that references a service device 110 or a specific service provided by the service device 110. In these cases, the script is executed by the client device 106 when the given electronic document is processed by the client device 106. The execution of the script configures the client device 106 to generate a request for a digital component 112 (referred to as a "component request"), which is transmitted to the service device 110 via the network 102. For example, the script may enable the client device 106 to generate a component request 112 that is packetized and includes header and payload data. The component request 112 may include event data specifying characteristics, such as the name (or network location) of the server from which the digital component is being requested, the name (or network location) of the requesting device (e.g., client device 106), and / or information that the service device 110 may use to select one or more digital components (e.g., target digital components 117 or other digital components from the database 116), or other content to be provided in response to the request. The component request 112 is transmitted by the client device 106 to a server of the service equipment 110 via the network 102 (eg, a telecommunications network).

[0036] The component request 112 may include event data specifying other event characteristics, such as characteristics of the electronic document being requested and the location of the electronic document where the digital component may be presented. For example, event data may be provided to the service device 110 that specifies a reference (e.g., a uniform resource locator (URL)) to an electronic document (e.g., a web page) in which the digital component is to be presented, available locations of the electronic document that may be used to present the digital component, the size of the available locations, and / or the media types that are eligible for presentation at these locations. Similarly, event data specifying keywords associated with the electronic document ("document keywords") or entities referenced by the electronic document (e.g., people, places, or things) may also be included in the component request 112 (e.g., as payload data) and provided to the service device 110 to facilitate identification of digital components that are eligible for presentation with the electronic document. The event data may also include a search query submitted from the client device 106 to obtain a search results page.

[0037] The component request 112 may also include event data related to other information, such as information that has been provided by a user of the client device, geographic information indicating the state or region in which the component request was submitted, or other information providing context for the environment in which the digital component will be displayed (e.g., the time of day of the component request, the day of the week of the component request, and the type of device in which the digital component will be displayed (such as a mobile device or a tablet device)). The component request 112 may be transmitted, for example, over a packetized network.

[0038] The service device 110 includes an artificial intelligence system 160 that implements one or more language model neural networks 170 (also referred to as "language models"), which may include large language models. Large language models ("LLMs") are models that are trained to generate and understand human language and / or computer code. LLMs are trained on large data sets of text and / or code, and they can be used for a variety of tasks. For example, LLMs can be trained to convert text from one language to another; summarize text, such as website content, search results, news articles, or research papers; answer questions about text, such as "What is the capital of Georgia?"; create chatbots that can converse with humans; and generate creative text, such as poetry and stories (in natural language) and computer code (in programming languages).

[0039] In some implementations, language model 170 is pre-trained, i.e., trained on language modeling tasks that do not require providing evidence in response to user questions, and service device 110 (e.g., using AI system 160) causes language model 170 to generate an output sequence according to a predetermined syntax through natural language prompts in the input sequence.

[0040] For example, the service device 110 (e.g., AI system 160) or a separate training system pre-trains the language model 170 (e.g., a neural network) on a language modeling task, such as a task that requires predicting the next word after the current sequence in the training data given a current text word sequence. As a specific example, the language model 170 can be pre-trained on a large text (e.g., text publicly available from the Internet or another text corpus) data set according to a maximum likelihood objective.

[0041] The language model 170 can be configured to perform any kind of language modeling task through training, i.e., can be configured to receive any kind of input prompt 172 (also referred to as “prompt” for short) and generate any kind of output sequence 174 (also referred to as “output” for short) based on the input 172. In general, the AI ​​system 160 receives the prompt 172 submitted to the language model 170 and causes the language model 170 to generate the output 174 as a response to the prompt 172.

[0042] As a specific example, language model 170 may be configured to perform a summarization task. In this example, AI system 160 receives prompt 172 from a user of client device 106, e.g., as part of or in addition to component request 112 or another request for output 174. Prompt 172 includes or references a set of online sources, and optionally includes or references instructions requiring AI system 160 to generate a summary of the set of online sources. To initiate creation of output 174, AI system 160 submits the received prompt 172 to one or more language models 170, which use prompt 172 to evaluate the information found at the online sources specified in prompt 172, and generate output 174 that summarizes the information.

[0043] The service device 110 selects a digital component (e.g., third-party content that may all be in the form of advertising content or non-advertising content, such as video files, audio files, images, text, game content, augmented reality content, and combinations thereof) to be presented with the output 174 as a response to the prompt 172, such as the target digital component 117.

[0044] In some implementations, the digital component is selected within a time that is less than a time threshold to avoid errors that may be caused by delayed selection of the digital component. For example, a delay in providing the digital component in response to the component request 112 may cause page load errors at the client device 106, or cause portions of the electronic document to remain unpopulated even after other portions of the electronic document are rendered at the client device 106.

[0045] As another example, as the latency in providing the digital component to the client device 106 increases, the topic to which the digital component relates is likely no longer relevant to the current conversational turn between the user and the language model 170 when the digital component is delivered to the client device 106. As a result, the user's experience with the AI ​​system 160 is negatively impacted.

[0046] Further, a delay in providing the digital component may result in a failure in the delivery of the digital component. For example, if the session is no longer ongoing at the client device 106, and the prompt 172 has become outdated when the digital component is provided.

[0047] In some implementations, the service device 110 is implemented in a distributed computing system that includes, for example, a server and a set of multiple computing devices that can operate together to perform operations of the language model 170. The set of multiple computing devices can also operate together to use the following reference Figures 2 to 5 Described techniques for responding to prompt 172 from a plurality of available digital components (DC 1-x )'s corpus that identifies and distributes the set of target digital components eligible for presentation with output 174.

[0048] In particular, a plurality of available digital components are stored in a database, for example in Figure 1 The digital component database 116 of FIG. 116 is indexed into different stages within a corresponding user journey for performing a computer-implemented task. A specific example of a user journey will be described in Figure 2 In short, each user journey represents a process during which a user performs actions (or operations) to complete the final goal of a task; each user journey is divided into an ordered sequence of stages during which different actions are performed by the user.

[0049] like Figure 1 As shown, the digital component database 116 stores these available digital components in the form of key-value pairs, where the keys represent different stages of the user journey and the values ​​represent the available digital components. Each key (stage) can reference the corresponding value (digital component) mapped to the key. The digital component database 116 can have any of a variety of known data structures to store keys and their associated values. These available digital components can be received by the service device 110 from one or more digital component providers (e.g., third-party content providers), for example, from multiple digital component providers corresponding to multiple different stages, respectively.

[0050] As shown in the figure, the digital component DC 1-2 Stored in association with the first stage S1, the digital component DC 100-101 Stored in association with the second stage S2, the digital component DC x-x+1 With the third stage S x The digital component database 116 thus stores a mapping between each of the different stages within the user journey and the corresponding digital component associated with the stage. Each stage within the user journey can be mapped to the same or different number (or category) of digital components. For example, different stages can be mapped to non-overlapping or partially overlapping subsets of all available digital components stored in the digital component database 116.

[0051] Identification of eligible digital components to be presented with output 174 as a response to prompt 172 includes identifying, by service device 110 and based on prompt 172, a target stage within the user journey, and then selecting one or more digital components mapped to the target stage as eligible digital components for presentation. Identification of the target stage includes using intermediate embeddings generated by one or more intermediate neural network layers of language model 170 during processing of prompt 172, as will be further described below.

[0052] In general, the service device 110 may select any number of target digital components that map to the identified target stage. In some implementations, the service device 110 may then generate reply data 120 (e.g., digital data representing the response) and transmit the reply data over the network 102, which enables the client device 106 to integrate the set of target digital components into the output 174, so that the set of target digital components (e.g., target third-party content) and the content of the output 174 generated by the language model 170 are presented together at the display of the client device 106.

[0053] In some implementations, the client device 106 executes instructions included in the reply data 120 that configure the client device 106 and enable the client device to obtain the target set of digital components from one or more digital component servers 108. For example, the instructions in the reply data 120 may include a network location (e.g., a URL) and a script that causes the client device 106 to transmit a server request (SR) 121 to the digital component server 108 to obtain a given winning digital component from the digital component server 108. In response to the request, the digital component server 108 will identify the given target digital component specified in the server request 121 (e.g., within a database 116 storing a plurality of digital components) and transmit digital component data (DC data) 122 to the client device 106 that presents the given target digital component at the client device 106 along with the output 174.

[0054] When the client device 106 receives the digital component data 122, the client device will render the digital component (e.g., third-party content) and present the digital component in an appropriate location. For example, the script 154 can create a walled garden environment, such as a frame, that is presented within the native content 152 of the electronic document 150, e.g., next to the native content. In some implementations, the digital component is overlaid on (or adjacent to) a portion of the native content 152 of the electronic document 150, and the service device 110 can specify the presentation location within the electronic document 150 in the reply 120. For example, when the native content 152 includes video content, the service device 110 can specify a location or object within a scene depicted in the video content above which the digital component is to be presented.

[0055] Figure 2is an example illustration of a user journey for performing a computer-implemented task according to an implementation of the present disclosure. A computer-implemented task may be any of a variety of tasks that require a user to perform an action using one or more user computing devices (e.g., client device 106). These actions may be performed in one or more software applications installed on the user computing devices, each with a corresponding user interface. For example, these actions may include selecting inputs by, for example, touch, gesture, or click, providing text, audio, or image input, and the like.

[0056] As a specific example, a computer-implemented task may involve identifying certain resources as target resources from among various resources available on the Internet, such as image files, audio files, video files, and web pages, and in some cases, utilizing the identified target resources to generate one or more solutions to a problem present in, for example, a real-world environment or a computer environment.

[0057] In this example, to complete such a computer-implemented task, the user typically navigates through (e.g., selects) multiple web pages to obtain information of interest. During the navigation, the user can submit prompts 172 to the AI ​​system 160 and correspondingly receive outputs 174 generated by the language model 170, which include or otherwise characterize resources. Additionally or alternatively, the user can submit queries to a search engine accessible by the service device 110, which can provide information about resources in a manner useful to the user.

[0058] The user typically performs different actions at each of the multiple stages included in the user journey. For example, at different stages of the user journey, the user selects different content presented on the user computing device, provides different text, audio or image input, or some combination of these.

[0059] In some implementations, Figure 2 The computer-implemented task in the example of can be a task of identifying a landing webpage representing a solution to a user's problem from a plurality of webpages. For example, the landing webpage can describe an item including, but not limited to, a commodity, product, service, experience, etc. that can be used as a solution to the problem.

[0060] like Figure 2As shown, the user journey includes a progression of five sequential stages. The beginning stage is the problem awareness stage 210. The problem awareness stage 210 may begin with the user viewing a source web page in a web browser of the client device 106. At the problem awareness stage 210, the user becomes aware of a particular problem that exists. Logically, the user may submit a prompt 172 to the language model 170 that defines the particular problem, thereby seeking answers about various ways to solve the particular problem. In response, the language model 170 returns information related to the prompt 172 as output 174 in a manner that reduces the amount of time the user spends sifting through various resources (e.g., various image files, audio files, video files, and web pages) to determine the information being sought.

[0061] The user journey proceeds from the problem awareness stage 210 to the solution provider awareness stage 220. The solution provider awareness stage 220 may begin after the user becomes aware of various possible solutions to a particular problem. At the problem awareness stage 210, the user may submit a prompt 172 describing a possible solution to the problem to the language model 170, thereby seeking information about existing solution providers that provide such solutions, and in response, receive output 174 from the language model 170 including information related to the prompt 172.

[0062] The user journey proceeds from the solution provider awareness stage 220 to the solution consideration stage 230. The solution consideration stage 230 may begin after the user becomes aware of existing solution providers that provide possible solutions to a particular problem. At the solution consideration stage 230, the user may submit a prompt 172 describing the existing solution providers to the language model 170, thereby seeking information about which specific problem solution(s) provided by one or more of the existing solution providers meet the needs of the user facing the problem, and in response, receive an output 174 from the language model 170 including information related to the prompt 172.

[0063] The user journey proceeds from the solution consideration stage 230 to the solution comparison stage 240. The solution comparison stage 240 may begin after the user becomes aware of a specific problem solution that meets the needs of the user facing the problem. At the solution comparison stage 240, the user may submit a prompt 172 describing a specific problem solution to the language model 170, thereby seeking information about which specific solution provider can provide the specific problem solution, and in response, receive an output 174 from the language model 170 including information related to the prompt 172.

[0064] The user journey proceeds from the solution comparison phase 240 to the final phase of the solution implementation phase 250. The solution implementation phase 250 may begin after the user has identified a particular solution provider, such as after the user arrives at a landing webpage corresponding to the particular solution provider in a web browser of the client device 106. At the solution implementation phase 250, the user may submit a prompt 172 to the language model 170 describing a particular problem solution provided by the identified particular solution provider, thereby seeking information about possible ways to implement the particular problem solution and / or what other solutions are needed to solve the particular problem, and in response, receive output 174 from the language model 170 including information related to the prompt 172.

[0065] A specific example of a user journey for completing a task is now described. It should be understood that alternative user journeys can complete the same task, and different users can perform different actions during alternative user journeys for completing the same task (but the stages included in those alternative user journeys for completing the same task will generally be similar to each other, e.g., the alternative user journeys will include the same number of stages).

[0066] Figure 3 Example operations for selecting a target digital component for presentation according to an implementation of the present disclosure are shown. These operations may be performed by Figure 1 The service device 110 is executed by the service device 110. The service device 110 includes an AI system 160, which in turn includes a language model 170.

[0067] AI system 160 receives prompts 172 and, in response, generates one or more outputs 174 conditioned on prompts 172 using language model 170. Language model 170 may be any suitable language model neural network that receives prompts 172 as an input sequence of text tokens selected from a vocabulary and autoregressively generates outputs 174 as an output sequence of text tokens from the vocabulary. For example, language model 170 may be a Transformer-based language model neural network or a recurrent neural network-based language model.

[0068] In some cases, language model 170 may be referred to as an autoregressive neural network when the neural network used to implement language model 170 autoregressively generates an output sequence of tokens. More specifically, the autoregressively generated output is created by generating each particular token in the output sequence conditioned on the current input sequence, including any tokens that precede the particular text token in the output sequence (i.e., tokens that have been generated for any previous positions in the output sequence that precede the particular position of the particular token), and contextual input that provides context for the output sequence.

[0069] For example, when generating a word-gram at any given position in the output sequence, the current input sequence may include the word-gram at any previous position in the input sequence and the output sequence that precedes the given position. As a specific example, the current input sequence may include the input sequence followed by the word-gram at any previous position in the output sequence that precedes the given position. Optionally, the input sequence and the current output sequence may be separated by one or more predetermined word-grams within the current input sequence.

[0070] More specifically, in order to generate a specific word-gram at a specific position in the output sequence, the neural network of the language model 170 may process the current input sequence to generate a score distribution (e.g., a probability distribution) that assigns a corresponding score (e.g., a corresponding probability) to each word-gram in the word-gram vocabulary. The neural network of the language model 170 may then use the score distribution to select a word-gram from the vocabulary as the specific word-gram. For example, the neural network of the language model 170 may greedily select the word-gram with the highest score, or may sample a word-gram from the distribution, for example, using kernel sampling or another sampling technique.

[0071] As a specific example, the language model 170 can be an autoregressive Transformer-based neural network that includes (i) multiple attention blocks, each of which applies a self-attention operation; and (ii) an output subnetwork that processes the output of the last attention block to generate a score distribution.

[0072] The language model 170 can have any of a variety of Transformer-based neural network architectures. Examples of such architectures include those described in the following literature: J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, DdL Casas, LA Hendricks, J. Welbl, A. Clark, et al. Training compute-optimal large language models, arXiv preprint arXiv:2203.15556, 2022; J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, H. F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche,LAHendricks,M.Rauh,P.Huang,A.Glaese,J.Welbl,S.Dathathri,S.Huang,J.Uesato,J.Mellor,I.Higg ins,A.Creswell,N.McAleese,A.Wu,E.Elsen,SMJayakumar,E.Buchatskaya,D.Budden,E.Sutherland,K.Simonyan, M.Paganini,L.Sifre,L.Martens,XLLi,A.Kuncoro,A.Nemazadeh,E.Gribovskaya,D.Donato,A.Lazaridou,A.Mens ch,J.Lespiau,M.Tsimpoukelli,N.Grigorev,D.Fritz,T.Sottiaux,M.Pajarskas,T.Pohlen,Z.Gong,D.Toyama,C.de Masson d'Autume,Y.Li,T.Terzi,V.Mikulik,I.Babuschkin,A.Clark,D.de Las Casas,A.Guy,C.Jones,J.Bradbury,M.Johnson,BAHechtman,L.Weidinger,I.Gabriel,WSIsaac,E.Lockhart,S.Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving. Scaling language models: Methods, analysis & insights from training gopher. CoRR, abs / 2112.11446, 2021; Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R.So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V.Le. Towards a human-like open-domain chatbot. CoRR, abs / 2001.09977, 2020; and Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint ar Xiv:2005.14165, 2020. .

[0073] like Figure 3 As shown, the Transformer-based neural network includes a sequence of attention blocks (e.g., attention blocks AC 180A-C), and during processing of a given input sequence, each attention block in the sequence receives a corresponding input hidden state for each input token in the given input sequence. The attention block then updates each of the hidden states at least in part by applying self-attention to generate a corresponding output hidden state for each of the input tokens. The input hidden state of the first attention block is an embedding of the input token in the input sequence, and the input hidden state of each subsequent attention block is an output hidden state generated by the previous attention block.

[0074] In this example, the output subnetwork processes the output hidden state generated by the last attention block in the sequence for the last input token in the input sequence to generate a score distribution.

[0075] In general, because language models are autoregressive, the service device 110 can use the same language model 170 to generate multiple different candidate output sequences in response to the same request, for example, by using beam search decoding from a score distribution generated by the language model 170, using a sampling and sorting decoding strategy, using different random seeds for a pseudo-random number generator used in sampling different runs of the language model 170, or using another decoding strategy that exploits the autoregressive properties of the language model.

[0076] To identify a target stage from among multiple stages within the user journey, the service device 110 uses output hidden states generated by one or more attention blocks during processing of the prompt 172 to generate the output 174 .

[0077] That is, the service device 110 uses an intermediate embedding in the latent space generated by one of the attention blocks based on the prompt 172 (e.g., the intermediate embedding generated by attention block B 180B), and then identifies one or more reference embeddings in the latent space based on corresponding distances between the one or more reference embeddings in the latent space and the intermediate embedding.

[0078] For example, the intermediate embedding may include a plurality of output hidden states that have been generated by the attention block for the input tokens (at least some of the input tokens) in prompt 172. As used in this specification, a "hidden state" or "embedding" is a vector of numerical values ​​(e.g., floating point values ​​or other values) having a predetermined dimension. The space of possible vectors having a predetermined number of dimensions is called a "latent space."

[0079] However, more generally, the intermediate embedding may include output activations generated by neurons of any intermediate neural network layer of the language model 170 during processing of the prompt 172 to generate the output 174, or a combination of output activations generated by neurons of two or more intermediate layers of the language model 170. For example, instead of or in addition to being the final output of the attention block, the intermediate embedding may include the output of a self-attention layer included in the attention block, the output of a feed-forward layer, or both.

[0080] The reference embedding may be generated by the language model 170 or another language model neural network based on processing historical cues (e.g., Figure 2 Different prompts in the example user journey corresponding to the problem awareness stage 210, solution provider awareness stage 220, solution consideration stage 230, solution comparison stage 240 and solution implementation stage 250, respectively, are pre-generated.

[0081] In some cases, these historical prompts may be actual prompts submitted by different users at different stages of the alternative user journey while completing the computer-implemented task. In some other cases, these historical prompts may be simulated prompts generated as output by language model 170, for example, based on contextual prompts using the generation capabilities of language model 170. Contextual prompts may, for example, include information available on a landing page.

[0082] The embeddings generated from these historical prompts can capture the semantic and / or syntactic properties of the historical prompts, as well as the context of the stage during which the prompt was submitted. In some implementations, these embeddings can take the form of "reference" embeddings that represent previous knowledge of the historical prompts obtained from the historical prompts by the language model 170. In other words, these reference embeddings map or project the historical prompts into a latent space. These reference embeddings can then be used to identify the target stage to which the new prompt 172 belongs from among multiple stages within the user journey.

[0083] exist Figure 3 In the example of , language model 170 or another language model neural network has previously been used to generate various reference embeddings belonging to regions 310, 320, and 330 of latent space 300. Each region corresponds to a stage in an alternative user journey. For example, region 310 may correspond to the first stage in the user journey. Clustering of reference embeddings generated based on processing similar prompts corresponding to the first stage ( Figure 3) may reside in a first region 310. Region 320 may correspond to a second stage of the user journey. Another cluster of reference embeddings generated based on processing similar prompts corresponding to the second stage may reside in a second region 320. And so on. These regions 310 to 330 may be defined in various ways, such as by using the largest bounding area or "convex hull" of all existing / known reference embeddings for a certain stage.

[0084] For the intermediate embedding 312 generated according to the prompt 172, in order to determine which stage among the multiple stages in the user journey the intermediate embedding 312 belongs to, one or more reference embeddings ( Figure 3 The one or more reference embeddings are identified by one or more distances between the intermediate embedding 312 in the latent space 300 and the intermediate embedding 312. These distances or "similarity" can be calculated in various ways, such as using cosine similarity, dot product, etc. For example, the reference embedding that is closest to the intermediate embedding can be identified and used to determine the corresponding target stage.

[0085] exist Figure 3 In the example of , the intermediate embedding 312 is closest to the reference embedding 314 in the latent space 300. The reference embedding 314 resides in the first region 310, which includes the cluster of reference embeddings generated based on processing the prompt corresponding to the first stage within the alternative user journey. Figure 3 In the example of , service device 110 determines that intermediate embedding 312 maps to the first region, and correspondingly determines that prompt 172 based on which intermediate embedding 312 has been generated maps to the first stage within the user journey.

[0086] Once the target stage within the user journey has been determined, the service device 110 may proceed to select one or more target digital components 117 from the available digital components mapped to the target stage. Figure 3 Selection of one or more target digital components 117 from the digital component subset 118A is shown to be delivered with the output 174 generated by the language model 170 and presented on the client device 106 .

[0087] In particular, the service device 110 may select any number of target digital components 117 from the digital component subset 118A, which is a specific subset of digital components that is mapped to the first stage within the user journey from among the plurality of different digital component subsets 118A-N stored in the database 116. For example, the database 116 may store each digital component in association with a distribution parameter that facilitates (e.g., triggers, regulates, or limits) the distribution / transmission of the corresponding digital component, and the service device 110 may then select as the target digital component one or more digital components having a distribution parameter that matches (e.g., matches exactly or matches with some pre-specified level of similarity) at least one criterion specified by the prompt 172 and / or the request for output 174 or otherwise derived from the prompt and / or the request for output.

[0088] Figure 4 is a flow chart of an example process for selecting a target digital component and presenting the target digital component together with a response to a natural language prompt according to an implementation of the present disclosure. The operations of process 400 may be performed, for example, by Figure 1 The process 400 may be performed by the service device 110 or another data processing device. The operation of process 400 may also be implemented as instructions stored on a computer-readable medium, which may be non-transitory. One or more data processing devices execute the instructions to cause one or more data processing devices to perform the operation of process 400.

[0089] A prompt is received from a user (402). The service device may, for example, receive a prompt that is typed and submitted by a user through a client device. The prompt typically includes natural language text, i.e., includes a plurality of input tokens included in a text token vocabulary, which includes, for example, one or more of characters, subwords, words, punctuation marks, numbers, or other symbols that appear in the natural language text.

[0090] The language model processes the prompt to generate a response to the prompt (404). The language model is typically trained on a large text and / or code data set to generate and understand human language and / or computer code. The response will similarly include natural language text, i.e., include a plurality of output tokens included in the text token vocabulary.

[0091] In some implementations, the language model is a Transformer-based neural network that includes a sequence of attention blocks, i.e., a plurality of attention blocks arranged in a sequence, wherein the output of any block except the last block is the input of another block in the block. During processing of a prompt, each attention block in the sequence receives a corresponding input hidden state for each input token in the prompt. The attention block then updates each of the hidden states at least in part by applying self-attention to generate a corresponding output hidden state for each of the input tokens.

[0092] An intermediate embedding is obtained (406). The intermediate embedding includes output activations generated by neurons of one or more intermediate layers of the language model during processing of the natural language prompt to generate a response. For example, the intermediate embedding may include a plurality of output hidden states that have been generated by one of the attention blocks for input tokens (at least some of the input tokens) in the prompt.

[0093] One or more target stages from among a plurality of different stages included in the user journey are determined based on the intermediate embedding (408). As discussed above, a user journey generally represents a process during which a user performs different actions to perform a computer-implemented task. Figure 5 Explaining step 408 in more detail, the figure shows sub-steps 502 to 508 corresponding to step 408 .

[0094] Go to Figure 5 , Figure 5 According to the implementation method of the present disclosure Figure 4 A flowchart of a process step with sub-steps within a step.

[0095] A distance is calculated between the intermediate embedding and each of the plurality of reference embeddings in the latent space (502). In some cases, the plurality of reference embeddings are generated by a language model based on processing historical prompts provided by different users at different reference stages included in the respective reference user journeys. In some other cases, the plurality of reference embeddings are generated by the language model based on processing simulated prompts, which are simulations of historical prompts that have been provided by different users at different reference stages included in the respective reference user journeys. Such simulated prompts can be generated, for example, by using a language model (leveraging its generative capabilities).

[0096] One or more reference embeddings having distances in the latent space that satisfy a distance threshold are selected (504). For example, the serving device may select one or more reference embeddings that are closest to the intermediate embedding (i.e., have the shortest distance to the intermediate embedding) from among all of the multiple reference embeddings. As another example, the serving device may select one or more reference embeddings that have distances relative to the intermediate embedding that are each below a threshold distance value. These distances may be calculated in a variety of ways, such as using cosine similarity, dot products, etc. As yet another example, the serving device may select one or more reference embeddings that are included in the same predefined region in the latent space as the intermediate embedding.

[0097] One or more reference stages within the corresponding reference user journey are identified (506). That is, the service device identifies one or more historical prompts, and the one or more reference embeddings selected at step 506 are generated by the language model based on the one or more historical prompts; and correspondingly determines one or more stages within the corresponding reference user journey during which the identified historical prompts are received by the language model as reference stages. The service device can select a corresponding stage for each historical prompt, for example, two or more different stages can be selected when multiple historical prompts are identified.

[0098] One or more reference phases are used as one or more target phases (508).

[0099] Back to Figure 4 , one or more target digital components are selected from the corresponding candidate digital components mapped to the one or more target phases (410). In some implementations, the available digital components are stored in a database in association with corresponding indexes representing different phases, and the service device can select any number of the available digital components associated with the index representing the reference phase as the one or more target digital components, for example by sampling the digital components with uniform randomness, or by selecting digital components with distribution parameters matching some criteria specified by the prompt or otherwise derived from the prompt.

[0100] One or more target digital components are presented for display to the user along with responses to the prompts that have been generated by the language model (412). For example, the service device may deliver both the target digital components and the responses to the client device, which then presents them on a computer display of the client device.

[0101] Figure 66 is a block diagram of an example computer system 600 that can be used to perform the operations described above according to implementations of the present disclosure. System 600 includes a processor 610, a memory 620, a storage device 630, and an input / output device 640. Each of components 610, 620, 630, and 640 can be interconnected, for example, using a system bus 650. Processor 610 is capable of processing instructions for execution within system 600. In one implementation, processor 610 is a single-threaded processor. In another implementation, processor 610 is a multi-threaded processor. Processor 610 is capable of processing instructions stored in memory 620 or on storage device 630.

[0102] The memory 620 stores information within the system 600. In one implementation, the memory 620 is a computer-readable medium. In one implementation, the memory 620 is a volatile memory unit. In another implementation, the memory 620 is a non-volatile memory unit.

[0103] The storage device 630 can provide mass storage for the system 600. In one implementation, the storage device 630 is a computer-readable medium. In various implementations, the storage device 630 may include, for example, a hard disk device, an optical disk device, a storage device shared by multiple computing devices over a network (e.g., a cloud storage device), or some other mass storage device.

[0104] The input / output device 640 provides input / output operations for the system 600. In one implementation, the input / output device 640 may include one or more of a network interface device (e.g., an Ethernet card), a serial communication device (e.g., an RS-232 port), and / or a wireless interface device (e.g., an 802.11 card). In another implementation, the input / output device may include a driver device configured to receive input data and send output data to other devices (e.g., a keyboard, a printer, a display, and other peripheral devices 660). However, other implementations may also be used, such as mobile computing devices, mobile communication devices, set-top TV client devices, etc.

[0105] Although in Figure 6 An example processing system is described in the specification, but the subject matter and implementation of the functional operations described in this specification may be implemented in other types of digital electronic circuit systems or in computer software, firmware or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them.

[0106] An electronic document (referred to simply as a document for brevity) need not correspond to a file. A document may be stored in a portion of a file that holds other documents, in a single file dedicated to the document in question, or in multiple coordinated files.

[0107] With respect to situations in which the systems discussed herein collect and / or use personal information about a user, the user may be provided with the opportunity to enable / disable or otherwise control programs or features that may collect and / or use personal information (e.g., information about the user's social network, social actions or activities, the user's preferences, or the user's current location). Additionally, certain data may be disposed of in one or more ways prior to its storage or use such that personally identifiable information associated with the user is deleted. For example, the user's identity may be anonymized such that personally identifiable information about the user cannot be determined, or the user's geographic location may be generalized (such as to a city, zip code, or state level) if location information is available such that the user's specific location cannot be determined.

[0108] The subject matter and the embodiments of the operation described in this specification may be implemented in a digital electronic circuit system or in computer software, firmware or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them. The embodiments of the subject matter described in this specification may be implemented as one or more computer programs, that is, one or more modules of computer program instructions, which are encoded on a computer storage medium for data processing equipment to execute or control the operation of the data processing equipment. Alternatively or in addition, the program instructions may be encoded on an artificially generated propagation signal (e.g., a machine-generated electrical, optical or electromagnetic signal), which is generated to encode information for transmission to a suitable receiver device for data processing equipment to execute. The computer storage medium may be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them, or included in them. In addition, although the computer storage medium is not a propagation signal, the computer storage medium may be the source or destination of the computer program instructions encoded in the artificially generated propagation signal. The computer storage medium may also be one or more separate physical components or media (e.g., multiple CDs, disks or other storage devices) or include them.

[0109] The operations described in this specification may be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0110] The term "data processing device" includes all types of devices, apparatuses and machines for processing data, including, for example, a programmable processor, a computer, a system on a chip, or a plurality of the foregoing or a combination thereof. The device may include a dedicated logic circuit system, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, the device may also include code that creates an execution environment for the computer program in question, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The device and the execution environment may implement a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.

[0111] This document relates to service devices. As used herein, a service device is one or more data processing devices that perform operations to facilitate the distribution of content over a network. A service device is depicted as a single block in a block diagram. However, although a service device can be a single device or a single group of devices, the present disclosure contemplates that a service device can also be a group of devices, or even a plurality of different systems that communicate to provide various content to a client device. For example, a service device can encompass one or more of a search system, a video streaming service, an audio streaming service, an email service, a navigation service, an advertising service, a gaming service, or any other service.

[0112] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages; and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network.

[0113] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuit systems, and the device can also be implemented as a special purpose logic circuit system, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).

[0114] For example, processors suitable for executing computer programs include general and special purpose microprocessors, and any one or more processors of any type of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more large-capacity storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to receive data from it or transfer data to it or both. However, a computer does not have to have such devices. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0115] To provide interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including sound, voice, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on a user's client device in response to a request received from the web browser.

[0116] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes a back-end component, such as a data server, or includes a middleware component, such as an application server, or includes a front-end component, such as a client computer with a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), interconnected networks (e.g., the Internet), and peer-to-peer networks (e.g., dedicated peer-to-peer networks).

[0117] A computing system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises from computer programs running on respective computers and having a client-server relationship to each other. In some embodiments, the server transmits data (e.g., an HTML web page) to a client device (e.g., for the purpose of displaying data to a user interacting with the client device and receiving user input from the user). Data generated at the client device (e.g., the result of a user interaction) may be received from the client device at the server.

[0118] Although this specification contains many specific implementation details, these details should not be interpreted as limitations on the scope of any invention or the scope that may be claimed, but should be interpreted as descriptions of features peculiar to specific embodiments of specific inventions. Certain features described in the context of separate embodiments in this specification may also be implemented in combination in a single embodiment. On the contrary, the various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. In addition, although features may be described above as working in certain combinations and even initially claimed as such, in some cases, one or more features from the claimed combination may be deleted from the combination, and the claimed combination may involve a sub-combination or a change in the sub-combination.

[0119] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in a continuous order, or that all of the operations shown be performed, to achieve the desired result. In some cases, multitasking and parallel processing can be advantageous. In addition, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0120] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions set forth in the claims can be performed in a different order and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In some implementations, multitasking and parallel processing can be advantageous.

Claims

1. A method performed by one or more computers, comprising: receiving natural language prompts from the user; processing the natural language prompt by a language model neural network to generate a response to the natural language prompt; obtaining an intermediate embedding generated by the language model neural network during processing of the natural language prompt to generate the response; determining one or more target stages from among a plurality of different stages included in a user journey based on the intermediate embeddings, during which a user performs different actions to perform a computer-implemented task; selecting one or more target digital components from corresponding candidate digital components mapped to the one or more target stages; as well as The one or more target digital components are presented together with the response to the natural language prompt that has been generated by the language model neural network for display to the user.

2. The method of claim 1, wherein: The computer-implemented tasks include: Navigating from a source web page through a plurality of web pages to reach a landing web page, wherein the landing web page represents a solution to the problem represented by the source web page.

3. The method according to any one of claims 1 to 2, wherein: The plurality of different phases is a progression of different phases including a problem awareness phase, followed by a solution provider awareness phase, followed by a solution consideration phase, followed by a solution comparison phase, and followed by a solution implementation phase.

4. The method according to any one of claims 1 to 3, wherein: The intermediate embedding includes: The output of the intermediate neural network layer of the language model neural network.

5. The method according to any one of claims 1 to 4, wherein: Determining the one or more target stages from among the plurality of different stages included in the user journey includes: calculating a distance between the intermediate embedding and each of a plurality of reference embeddings in a latent space, wherein the plurality of reference embeddings are generated by the language model neural network based on processing historical natural language prompts provided by different users at different reference stages included in corresponding reference user journeys; selecting one or more reference embeddings having distances satisfying a distance threshold; identifying one or more reference stages within the respective reference user journeys during which the one or more reference embeddings were generated by the language model neural network; and The one or more reference phases are used as the one or more target phases.

6. The method of claim 5, comprising: The language model neural network is used to generate additional historical natural language prompts based on information available on the landing web page, the additional historical natural language prompts representing the historical natural language prompts that would be provided by the different users at the different reference stages included in the corresponding reference user journeys.

7. The method according to any one of claims 1 to 6, wherein: Selecting the one or more target digital components comprises: A respective candidate digital component associated with each stage is received from a digital component provider and for each stage of the plurality of different stages included in the user journey.

8. A system comprising one or more computers and one or more storage devices storing instructions operable, when executed by the one or more computers, to cause the one or more computers to perform the operations of the corresponding method as claimed in any preceding claim.

9. A computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the corresponding method as recited in any preceding claim.