Generative artificial intelligence
The artificial intelligence system generates prompts and uses language models to generate clauses, combining post-processing operation evaluation and ranking candidate digital components, solves the problem of difficulty in generating creative and factual digital components in the existing technology, and achieves high-quality and efficient output.
Patent Information
- Application Number
- CN202480004340.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-22
- Filing Date
- 2024-05-22
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is difficult to effectively generate creative and factual digital components, especially under restrictive conditions, to ensure the quality and authenticity of outputs.
Generate prompts through artificial intelligence systems, generate clauses using language models, and evaluate and rank candidate digital components through post-processing operations, and ultimately supply high-quality output digital components. Generating prompts may include specifying the source, entity constraints, and by-demand constraints of the online content to ensure the accuracy and relevance of the content.
It realizes the generation of high-quality, creative and factual digital components under restrictive conditions, improves the utilization efficiency of computing resources, and quickly generates output in a real-time interactive environment.
Smart Images

Figure CN119968624A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Application No. 63 / 503,685, filed on May 22, 2023. The disclosure of the prior application is considered part of the disclosure of the present application and is incorporated by reference into the disclosure of the present application. Background Art
[0003] This specification relates to data processing and generative artificial intelligence.
[0004] Advances in machine learning have enabled artificial intelligence to be implemented in more applications. For example, large language models have been implemented to allow conversational interactions with computers using natural language rather than a restricted set of prompts. This allows for more natural interactions with computers. Summary of the invention
[0005] In general, one innovative aspect of the subject matter described in this specification can be embodied in a method comprising the following actions: generating, by an artificial intelligence system, a prompt comprising a query and a set of constraints that restrict clauses generated by a language model, wherein the set of constraints comprises a summary of a specified source of online content; generating, by the artificial intelligence system, a plurality of candidate digital components using the clauses generated by the language model using the summary of the specified source of online content; performing, by the artificial intelligence system, one or more post-processing operations that evaluate one or more characteristics of the plurality of candidate digital components; ranking, by the artificial intelligence system, each of the plurality of candidate digital components based on the post-processing operations; and supplying, by the artificial intelligence system, at least one output digital component in the set of highest-ranked candidate digital components.
[0006] These and other embodiments may each optionally include one or more of the following features. The method may include: collecting passages from a collection of online resources using a site-constrained query that requires passages to be collected from one or more network locations specified in a site constraint; and summarizing the passages collected from the one or more network locations into passage summaries.
[0007] Generating the prompt may include inserting at least a portion of the passage summary into the prompt as a context constraint that limits content created by the language model to topics specified in the context constraint.
[0008] Generating the hint may include inserting an entity name of an entity referenced by the one or more network locations into the hint as an entity constraint specifying that content identifying the entity must be included in the content created by the language model.
[0009] Generating the hint may include inserting a grounding constraint into the hint, the grounding constraint requiring that content created by the language model be present in a specified set of online resources.
[0010] Inserting the basis constraint into the hint may include inserting a second level domain into the hint, the second level domain requiring that content created by the language model be present in a resource within the second level domain.
[0011] Generating the plurality of candidate digital components may include combining an output of the language model with a link to the secondary domain.
[0012] Performing one or more post-processing operations to evaluate one or more characteristics of multiple candidate digital components may include: generating a basis score for each clause of the output of the language model, the basis score specifying the likelihood that the clause is factual; filtering the output of the language model by removing one or more clauses having a basis score that fails to meet a basis threshold, the basis threshold being demarcated between clauses classified as factual and non-factual; and replacing the one or more removed clauses with another clause of the output having a basis score that meets the basis threshold.
[0013] Performing one or more post-processing operations to evaluate one or more characteristics of multiple candidate digital components may include: for each given candidate digital component in the multiple candidate digital components: evaluating the relevance of clauses in the given candidate digital component to the prompted query; evaluating the completeness level of the clauses, the completeness level specifying how comprehensively the clauses in the given candidate digital component describe the topic in the secondary domain linked to by the given candidate digital component; and evaluating the tone of the clauses of the candidate digital component to determine whether the clauses characterize the item as positive or negative.
[0014] The details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a block diagram of an example environment in which generative artificial intelligence can be implemented.
[0016] Figure 2 is a block diagram illustrating the interaction between an artificial intelligence system, a language model, and a client device.
[0017] Figure 3 is a flow chart of an example process for AI to generate creative and factual digital components.
[0018] Figure 4is a block diagram of an example computer.
[0019] Like reference numbers and designations in the various drawings refer to like elements. DETAILED DESCRIPTION
[0020] This specification describes technologies that enable artificial intelligence to generate creative and factual new digital components. Artificial intelligence (AI) is a branch of computer science that focuses on creating intelligent agents that can learn and act autonomously (e.g., without human intervention). Artificial intelligence systems can utilize one or more of the following: (i) machine learning, which focuses on developing algorithms that can learn from data; (ii) natural language processing, which focuses on understanding and generating human language; and / or (iii) computer vision, which is a field focused on understanding and interpreting images and videos. Artificial intelligence systems can include generative models that generate new content (e.g., images / video, text, audio, or other content) in response to input prompts.
[0021] The techniques described throughout this specification enable artificial intelligence to generate a large number of new digital components using various combinations of text and / or images, which are not only creative but also factual. For example, an artificial intelligence system can collect information from various sources (such as various web pages or other trusted online resources) and combine this information in different ways to create different candidate digital components. Generally speaking, the system uses input prompts to a language model (such as a large language model (LLM)), which outputs multiple clauses. The system uses the clauses to create multiple different digital components, and then performs post-processing to select output digital components from the different candidate digital components.
[0022] As discussed in more detail below, the hints are specifically (e.g., created or enhanced) to improve the overall quality of the generated candidate digital components. The generated candidate digital components are then evaluated against each other using post-processing operations to determine which candidate digital components are of higher quality than other candidate digital components (e.g., given the current context), and one or more of the higher quality digital components are output to a computing device (e.g., a user computer, a mobile device, a tablet device, an audio device, a gaming device, etc.).
[0023] The use of specialized prompts reduces wasted computing resources that would otherwise generate more low-quality digital components if more general prompts were used. Similarly, as discussed in more detail below, by using specialized prompts to constrain the parameters used by the language model to generate candidate digital components, the number of generated candidate digital components can be reduced, thereby saving computing resources and generating output more quickly. For example, by constructing prompts to limit the types of content that can be included in the generated candidate digital components, the language model will not generate candidate digital components that violate the constraints in the prompts, thereby avoiding the creation of unwanted candidate digital components, which reduces the time required to generate candidate digital components, the memory required to store candidate digital components, and the computing resources required to generate and evaluate candidate digital components. This all facilitates the system to be able to create new digital components more quickly, so that the new digital components can be created and supplied in a real-time interactive environment (e.g., in response to a user search query).
[0024] The post-processing operation may include, for example, evaluating the candidate digital components based on various criteria, and scoring each of the candidate digital components based on the evaluation. For example, a post-processing operation may perform a prediction of the possibility that a specific candidate digital component is unfounded (e.g., including information that cannot be verified in a specified corpus). Using this type of post-processing operation allows for looser constraints in the construction of specialized prompts, which may allow the language model to generate more creative candidate digital components, while still ensuring that the output digital component has at least a baseline level of authenticity. The post-processing operation may also use various heuristics to evaluate the different characteristics of each of the candidate digital components, and may assign scores based on various heuristics. In some implementations, the scores are weighted and aggregated to produce a final score, which is used to rank the candidate digital components. Additionally or alternatively, a machine learning model may be trained to score the quality of the digital components, and those scores may be used to rank the candidate digital components. One or more of the highest-ranked candidate digital components are then selected as output digital components to be supplied.
[0025] As used throughout this document, the phrase "digital component" refers to a discrete unit of digital content or digital information (e.g., a video clip, an audio clip, a multimedia clip, game content, an image, text, a bullet point, an artificial intelligence output, a language model output, or another unit of content). A digital component may be electronically stored in a physical memory device as a single file or as a collection of files, and a digital component may take the form of a video file, an audio file, a multimedia file, an image file, or a text file, and include advertising information, such that an advertisement is a type of digital component.
[0026] Figure 11 is a block diagram of an example environment 100 in which generative artificial intelligence can be implemented. The example environment 100 includes a network 102, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof. The network 102 connects electronic document servers 104, user devices 106, digital component servers 108, and service devices 110. The example environment 100 can include many different electronic document servers 104, user devices 106, and digital component servers 108.
[0027] Client devices 106 are electronic devices capable of requesting and receiving online resources over network 102. Example client devices 106 include personal computers, gaming devices, mobile communication devices, digital assistant devices, augmented reality devices, virtual reality devices, and other devices that can send and receive data over network 102. Client devices 106 typically include user applications (such as a web browser) to facilitate sending and receiving data over network 102, but native applications executed by client devices 106 may also facilitate sending and receiving data over network 102.
[0028] Game device is the device that enables the user to participate in game application, for example, in the device, the user can control one or more roles, avatars or other rendered contents presented in the game application.Game device typically includes a computer processor, a memory device and a controller interface (physical or visually rendered) that enables the user to control the contents rendered by the game application.Game device can store and execute game application locally, or execute the game application (for example, online game application) stored and / or served at least in part by a cloud server.Similarly, game device can be docked with a game server that executes game application and "streams" the game application to game device.Game device can be a tablet device, a mobile telecommunication device, a computer or another device that also performs other functions in addition to executing the game application.
[0029] The digital assistant device includes a device with a microphone and a speaker. The digital assistant device is generally capable of receiving input by voice, and responding to the content using audible feedback, and can present other audible information. In some cases, the digital assistant device also includes a visual display or communicates with a visual display (e.g., by wireless or wired connection). When there is a visual display, feedback or other information can also be provided visually. In some cases, the digital assistant device can also control other devices, such as lights, locks, cameras, climate control devices, alarm systems, and other devices registered with the digital assistant device.
[0030] As shown, client device 106 is presenting electronic document 150. An electronic document is data that presents a collection of content at client device 106. Examples of electronic documents include web pages, word processing documents, portable document format (PDF) documents, images, videos, search result pages, and feeds. Native applications (e.g., "apps" and / or game applications) such as applications installed on mobile, tablet, or desktop computing devices are also examples of electronic documents. Electronic documents can be provided to client devices 106 by electronic document servers 104 ("Electronic Doc Servers").
[0031] For example, the electronic document server 104 may include a server hosting a publisher website. In this example, the client device 106 may initiate a request for a given publisher web page, and the electronic server 104 hosting the given publisher web page may respond to the request by sending a machine executable instruction that initiates presentation of the given web page at the client device 106.
[0032] In another example, the electronic document server 104 may include an app server from which the client device 106 may download the app. In this example, the client device 106 may download the files required to install the app at the client device 106, and then execute the downloaded app locally (i.e., on the client device). Alternatively or additionally, the client device 106 may initiate a request to execute the app, which is sent to the cloud server. In response to receiving the request, the cloud server may execute the application and stream the user interface of the application to the client device 106, so that the client device 106 does not have to execute the application app itself. Instead, the client device 106 may present the user interface generated by the cloud server executing the app, and communicate any user interaction with the user interface back to the cloud server for processing.
[0033] An electronic document may include a variety of content. For example, an electronic document 150 may include native content 152 that is within the electronic document 150 itself and / or does not change over time. An electronic document may also include dynamic content that may change over time or based on each request. For example, a publisher of a given electronic document (e.g., electronic document 150) may maintain data sources for populating various parts of the electronic document. In this example, a given electronic document may include a script, such as script 154, that causes a client device 106 to request content (e.g., digital components) from a data source when the given electronic document is processed (e.g., rendered or executed) by a client device 106 (or a cloud server). The client device 106 (or cloud server) integrates the content (e.g., digital components) obtained from the data source into a given electronic document to create a composite electronic document including the content obtained from the data source.
[0034] In some cases, a given electronic document (e.g., electronic document 150) may include a digital component script (e.g., script 154) that references a service device 110 or a specific service provided by the service device 110. In these cases, the digital component script is executed by the client device 106 when the given electronic document is processed by the client device 106. The execution of the digital component script configures the client device 106 to generate a request for a digital component 112 (referred to as a "component request"), which is sent to the service device 110 via the network 102. For example, the digital component script may enable the client device 106 to generate a packetized data request including header and payload data. The component request 112 may include event data specifying characteristics, such as the name (or network location) of the server from which the digital component is being requested, the name (or network location) of the requesting device (e.g., client device 106), and / or information that the service device 110 may use to select one or more digital components or other content to be provided in response to the request. The component request 112 is sent by the client device 106 to a server of the service device 110 via the network 102 (e.g., a telecommunications network).
[0035] The component request 112 may include event data specifying other event characteristics, such as characteristics of the electronic document being requested and the location of the electronic document at which the digital component may be presented. For example, event data specifying a reference (e.g., a URL) to an electronic document (e.g., a web page) in which the digital component is to be presented, an available location of the electronic document that may be used to present the digital component, the size of the available location, and / or a media type that is eligible for presentation at the location may be provided to the service device 110. Similarly, event data specifying keywords associated with the electronic document ("document keywords") or entities referenced by the electronic document (e.g., people, places, or things) may also be included in the component request 112 (e.g., as payload data) and provided to the service device 110 to facilitate identification of digital components eligible for presentation with the electronic document. The event data may also include a search query submitted from the client device 106 to obtain a search results page.
[0036] The component request 112 may also include event data related to other information, such as information that has been provided by a user of the client device, geographic information indicating the state or region from which the component request is being submitted, or other information providing the context of the environment in which the digital component will be displayed (e.g., the time of day of the component request, the day of the week of the component request, the type of device at which the digital component will be displayed, such as a mobile device or a tablet device). The component request 112 may be sent, for example, over a packetized network, and the component request 112 itself may be formatted as packetized data having a header and payload data. The header may specify a destination for the packet, and the payload data may include any of the information discussed above.
[0037] The service device 110 selects a digital component (e.g., third-party content such as video files, audio files, images, text, game content, augmented reality content, and combinations thereof, all of which may take the form of advertising content or non-advertising content) to be presented with a given electronic document (e.g., at a location specified by the script 154) in response to receiving the component request 112 and / or using information included in the component request 112.
[0038] In some implementations, the digital component is selected in less than one second to avoid errors that may result from delayed selection of the digital component. For example, a delay in providing the digital component in response to the component request 112 may cause page loading errors at the client device 106 or cause portions of the electronic document to remain unpopulated even after other portions of the electronic document are presented at the client device 106.
[0039] Moreover, as the delay in providing the digital component to the client device 106 increases, it becomes more likely that the electronic document will no longer be present at the client device 106 when the digital component is delivered to the client device 106, thereby negatively affecting the user's experience of the electronic document. Additionally, the delay in providing the digital component may cause the delivery of the digital component to fail, for example, if the electronic document is no longer present at the client device 106 when the digital component is provided.
[0040] In some implementations, the service device 110 is implemented in a distributed computing system that includes, for example, a server and a set of multiple computing devices 114 that are interconnected and identify and distribute digital components in response to the request 112. The set of multiple computing devices 114 operate together to select from millions of available digital components (DC 1-x ) to identify a set of digital components that are eligible for presentation in an electronic document. Millions of available digital components may be indexed, for example, in the digital component database 116. Each digital component index entry may reference a corresponding digital component and / or include distribution parameters (DP1 to DP2) that facilitate (e.g., trigger, constrain, or limit) the distribution / transmission of the corresponding digital component. x ). For example, the distribution parameters may facilitate (e.g., trigger) the transmission of the digital component by requiring the component request to include at least one criterion that matches (e.g., completely or with some pre-specified level of similarity) one of the distribution parameters of the digital component.
[0041] In some implementations, the distribution parameters for a particular digital component may include a distribution keyword that must be matched (e.g., to an electronic document, a document keyword, or a term specified in a component request 112) in order for the digital component to be eligible for presentation. Additionally or alternatively, the distribution parameters may include embeddings that may use a variety of different data dimensions, such as website details and / or consumption details (e.g., page viewports, user scrolling speed, or other information about data consumption). The distribution parameters may also require that the component request 112 include information specifying a particular geographic region (e.g., a country or state) and / or information specifying that the component request 112 originated from a particular type of client device (e.g., a mobile device or a tablet device) in order for the digital component to be eligible for presentation. The distribution parameters may also specify an eligibility value (e.g., a ranking score, or some other specified value) that is used to evaluate the eligibility of the digital component for distribution / transmission (e.g., with other available digital components).
[0042] The identification of eligible digital components may be split into a plurality of tasks 117a to 117c, which are then assigned among the computing devices within the set of the plurality of computing devices 114. For example, different computing devices 114 in the set may each analyze different portions of the digital component database 116 to identify various digital components having distribution parameters that match the information included in the component request 112. In some implementations, each given computing device 114 in the set may analyze different data dimensions (or sets of dimensions) and communicate (e.g., send) results of the analysis (Res 1 to Res 3) 118a to 118c back to the service device 110. For example, the results 118a to 118c provided by each of the computing devices 114 in the set may identify a subset of digital components that are eligible for distribution in response to the component request and / or a subset of digital components having certain distribution parameters. The identification of the subset of digital components may include, for example, comparing the event data to the distribution parameters, and identifying a subset of digital components having distribution parameters that match at least some features of the event data.
[0043] The service device 110 aggregates the results 118a-118c received from the set of multiple computing devices 114 and uses information associated with the aggregated results to select one or more digital components to be provided in response to the request 112. For example, the service device 110 can select the set of winning digital components (one or more digital components) based on the results of one or more content evaluation processes, as described below. In turn, the service device 110 can generate and send reply data 120 (e.g., digital data representing the reply) over the network 102, which enables the client device 106 to integrate the set of winning digital components into a given electronic document so that the set of winning digital components (e.g., winning third-party content) and the content of the electronic document are presented together at the display of the client device 106.
[0044] In some implementations, the client device 106 executes instructions included in the reply data 120 that configure the client device 106 and enable it to obtain the set of winning digital components from one or more digital component servers 108. For example, the instructions in the reply data 120 may include a network location (e.g., a uniform resource locator (URL)) and a script that causes the client device 106 to send a server request (SR) 121 to the digital component server 108 to obtain a given winning digital component from the digital component server 108. In response to the request, the digital component server 108 will identify the given winning digital component specified in the server request 121 (e.g., within a database storing a plurality of digital components) and send digital component data (DC data) 122 to the client device 106 that presents the given winning digital component in an electronic document at the client device 106.
[0045] When the client device 106 receives the digital component data 122, the client device will render the digital component (e.g., third-party content) and present the digital component at a location specified by or assigned to the script 154. For example, the script 154 may create a walled garden environment (such as a frame) that is presented within, e.g., next to, the native content 152 of the electronic document 150. In some implementations, the digital component is overlaid on (or adjacent to) a portion of the native content 152 of the electronic document 150, and the service device 110 may specify the presentation location within the electronic document 150 in the reply 120. For example, when the native content 152 includes video content, the service device 110 may specify a location or object within a scene depicted in the video content over which the digital component is to be presented.
[0046] The service device 110 may also include an artificial intelligence system 160 configured to autonomously generate a digital component prior to (e.g., offline) and / or in response to the request 112 (e.g., online or in real time). As described in more detail throughout this specification, the artificial intelligence ("AI") system 160 may collect online content about a particular entity (e.g., a digital component provider or another entity) and summarize the collected online content using one or more language models 170, which may include a large language model.
[0047] Large language models ("LLMs") are models trained to generate and understand human language. LLMs are trained on large datasets of text and code, and they can be used for a variety of tasks. For example, LLMs can be trained to translate text from one language to another; summarize text, such as website content, search results, news articles, or research papers; answer questions about text, such as "What is the capital of Georgia?"; create chatbots that can have conversations with humans; and generate creative text, such as poetry, stories, and code.
[0048] The language model 170 may be any suitable language model neural network that receives an input sequence consisting of text tokens selected from a vocabulary and autoregressively generates an output sequence consisting of text tokens from the vocabulary. For example, the language model 170 may be a Transformer-based language model neural network or a recurrent neural network-based language model.
[0049] In some cases, language model 170 may be referred to as an autoregressive neural network when the neural network used to implement language model 170 autoregressively generates an output sequence of tokens. More specifically, the autoregressively generated output is created by generating each particular token in the output sequence conditioned on a current input sequence, including any tokens preceding the particular text token in the output sequence (i.e., tokens that have been generated for any previous positions in the output sequence preceding the particular position of the particular token) and contextual input that provides context for the output sequence.
[0050] For example, when generating a word-gram at any given position in the output sequence, the current input sequence may include the input sequence and the word-gram at any preceding position in the output sequence before the given position. As a specific example, the current input sequence may include the input sequence followed by the word-gram at any preceding position in the output sequence before the given position. Optionally, the input and current output sequences may be separated by one or more predetermined word-grams within the current input sequence.
[0051] More specifically, to generate a particular word-gram at a particular position within the output sequence, the neural network of the language model 170 may process the current input sequence to generate a score distribution (e.g., a probability distribution) that assigns a corresponding score (e.g., a corresponding probability) to each word-gram in the vocabulary of word-grams. The neural network of the language model 170 may then use the score distribution to select a word-gram from the vocabulary as the particular word-gram. For example, the neural network of the language model 170 may greedily select the word-gram with the highest score or may sample word-grams according to the distribution, for example, using kernel sampling or another sampling technique.
[0052] As a specific example, the language model 170 can be an autoregressive transformer-based neural network that includes (i) multiple attention blocks, each of which applies a self-attention operation; and (ii) an output subnetwork that processes the output of the last attention block to generate a score distribution.
[0053] The language model 170 may have any of a variety of transformer-based neural network architectures. Examples of such architectures include those described in: J. Hoffmann, S. Borgeaud, A. Mensch, E.Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, LA Hendricks, J.Welbl, A. Clark, et al. Training compute-optimal large language models, arXiv preprint arXiv:2203.15556, 2022; J.W Rae, S.Borgeaud, T. Cai, K. Millican, J. Hoffmann, HF Song, J. Aslanides, S.Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A.Cassirer, R. Powell, G. van den Driessche, LA Hendricks, M. Rauh, P.Huang, A. Glaese, J. Welbl, S. Dathathri, S. Huang, J. Uesato, J. Mellor, I.Higgins, A. Creswell, N. McAleese, A.Wu, E. Elsen, SM Jayakumar, E.Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M. Paganini, L. Sifre, L.Martens, XL Li, A. Kuncoro, A. Nematzadeh, E. Gribovskaya, D. Donato, A.Lazaridou, A. Mensch, J. Lespiau, M. Tsimpoukelli, N. Grigorev, D. Fritz, T.Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D. Toyama, C. de Masson d'Autume,Y. Li, T. Terzi, V. Mikulik, I.Babuschkin, A. Clark, D. de Las Casas, A. Guy, C. Jones, J. Bradbury, M. Johnson, BA Hechtman, L. Weidinger, I. Gabriel, WS Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving. Scaling language models: Methods, analysis & insights from training gopher. CoRR, abs / 2112.11446, 2021; Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer (Exploring the Limits of Transfer Learning with the Unified Text-to-Text Transformer). arXiv preprint arXiv:1910.10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. Towards a human-like open-domain chatbot. CoRR, abs / 2001.09977, 2020; and Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, PranavShyam, Girish Sastry, Amanda Askell et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020. .
[0054] However, in general, a transformer-based neural network includes a sequence of attention blocks, and during processing of a given input sequence, each attention block in the sequence receives a respective input hidden state for each input token in the given input sequence. The attention blocks then update each of the hidden states at least in part by applying self-attention to generate a respective output hidden state for each of the input tokens. The input hidden state for the first attention block is an embedding of the input token in the input sequence, and the input hidden state for each subsequent attention block is the output hidden state generated by the previous attention block.
[0055] In this example, the output subnetwork processes the output hidden state generated by the last attention block in the sequence for the last input token in the input sequence to generate a score distribution.
[0056] In general, because language models are autoregressive, the service device 110 can use the same language model 170 to generate multiple different candidate output sequences in response to the same request, for example, by using beam search decoding based on the score distribution generated by the language model 170, using a sampling and ranking decoding strategy, by using different random seeds for the pseudo-random number generator used in sampling of different runs of the language model 170, or using another decoding strategy that exploits the autoregressive properties of the language model.
[0057] In some implementations, language model 170 is pre-trained, i.e., trained on language modeling tasks that do not require providing evidence in response to user questions, and service device 110 (e.g., using AI system 160) causes language model 170 to generate an output sequence according to a predetermined grammar via natural language prompts in an input sequence.
[0058] For example, service device 110 (e.g., AI system 160) or a separate training system pre-trains language model 170 (e.g., a neural network) on a language modeling task (e.g., a task requiring, given a current sequence of text tokens, to predict the next token in the training data that follows the current sequence). As a specific example, language model 170 may be pre-trained on a large dataset of text (e.g., text publicly available from the Internet or another text corpus) with a maximum likelihood objective.
[0059] In some implementations, AI system 160 may generate prompts 172 that are submitted to language model 170 and cause language model 170 to generate output sequence 174, also referred to as a passage or simply “output”. AI system 160 may generate prompts in a manner (e.g., having a certain structure) that identifies a list of online sources of information, such as a list of websites or data repositories, and specifies a set of constraints that language model 160 must use to generate a summary of the information found at the online sources specified in prompt 172. To initiate creation of output sequence 174, AI system 160 submits prompt 172 to one or more language models 170, which use prompt 172 to evaluate the information found at the online sources specified in prompt 172 and generate output 174 that summarizes the information according to the constraints specified in prompt 172.
[0060] The AI system 160 may use the generated summary as part of another prompt 172 sent to the language model 170. For example, the AI system 160 may insert the generated summary into an additional prompt 172 (e.g., a prompt generated after receiving the summary), which is submitted to the language model 170 as a constraint for generating a clause for use in a digital component being generated by the AI system 160. More specifically, assume that the AI system 160 is generating a digital component to be provided in response to a request 112 including a keyword / query. In this example, the AI system 160 may generate an additional prompt 172 to include a query and a set of constraints (including the summary received in the previous output 174). The set of constraints of the additional prompt 172 may also include instructions on how the clause generated by the language model 170 using the additional prompt 172 is to be formatted, styled, semantically styled, and other (e.g., specifying what should be excluded from the clause, such as granular details, such as numbers). For example, the additional prompt 172 may take the following form:
[0061] Write good_output that searches for the digital component where the query is "10g network" and the entity is "example_network_provider". good_output is based only on the key parts of this summary: "200Mbps internet with WiFi equipment on the example_network_provider 10G Network— $50 / mo for 2 years. example_network_provider 10G Network is getting fasterand more reliable every day. example_network_provider gives you all this andthen some. example_network_provider Mobile is the fastest mobile service andthe best price for 2 lines of Unlimited. Stream the latest season whereveryou go. Storm-Ready WiFi that's backed by a wireless connection, unlimiteddata, and a battery. (example_network_provider's 10G network provides 200 Mbps Internet and WiFi equipment—$50 / mo for 2 years. example_network_provider's 10G network is getting faster and more reliable every day. example_network_provider gives you all of this and more. example_network_provider mobile is the fastest mobile service and offers the best price for 2 unlimited lines. Stream the latest season of shows on the go. WiFi that's taking you by storm with wireless connectivity, unlimited data, and battery. )"good_output must be in bullet point format.good_output must have exactly 3 bullet points.Each bullet point must be less than 90 characters.good_output may not have nested bullet points.good_output must be compelling and demonstrate a value proposition.good_output must be useful and informative, and must avoid boring details (such as numbers).
[0062] In this example prompt, AI system 160 is providing the following constraints to language model 170:
[0063] - The query constraint specifies the query "10G Network" that the output clause should be related to.
[0064] - Entity constraint specifies "example_network_provider" as the entity name to be used in the output clause.
[0065] - A profile constraint specifies the content profile to be used during clause generation, i.e., "200 Mbps internetwith WiFi...".
[0066] - The "must be in bullet-point format ... must have exactly 3 bullet-points. Each bullet-point must be less than 90 characters ... must have nonested bullets" style constraint specifies the format that the output clause must use.
[0067] - The semantic / tone control constraint of “must be catchy and show value-prop. good_output must be useful and informative, and must avoid boring details like numbers” defines the tone and content of the output clause generated using the prompt.
[0068] Submitting this additional hint 172 to the language model 170 causes the language model to generate an additional output 174, which includes multiple sets of clauses generated from the query and the constraints, which is electronically communicated to the AI system 160. The AI system 160 receives the clauses of the additional output 174 and generates multiple candidate digital components that can be provided in response to the request 112. In some implementations, each different candidate digital component includes a different combination of the clauses received from the language model 170 in the additional output 174. For example, assuming that the additional output 174 includes 12 different clauses, and the formatting of the digital components being generated by the AI system 160 each includes space for three different clauses, the AI system 160 can use 3 different clauses in each of the candidate digital components to make 220 different candidate digital components (e.g., 12! / (3!(12-3)!)=220). In some cases, the AI system 160 may also create candidate digital components using a set of different links to online content (e.g., second-level domain links to web pages, phone numbers, etc. that discuss topics related to the candidate digital components), which may continue to exponentially increase the number of different candidate digital components that the AI system 160 may create using clauses of the additional output 174 of the language model 170.
[0069] The AI system 160 may perform one or more post-processing operations to evaluate one or more characteristics of the plurality of candidate digital components. In some implementations, the post-processing operations may include generating a basis score for each clause from the additional output 174. The basis score is a value that specifies the likelihood that the clause is factual. Figure 2 The generation of the basis scores is discussed in more detail.
[0070] Post-processing operations may also include an evaluation of the relevance of the clause to the query constraints, the level of completeness of the clause relative to the content at the link included in the candidate digital component, and / or an evaluation of the tone (e.g., positive or negative) of the clause. Figure 2 As discussed in more detail, post-processing operations may be used to score or otherwise assign a priority level to each of the candidate digital components, such that the AI system 160 may rank multiple candidate digital components relative to each other and ultimately supply one or more of the highest ranked candidate digital components as output digital components in reply 120 to the request 112. It should be noted that while the operations of the AI system 160 and the language model 170 are described above as being performed in response to receiving the request 112, at least some of the operations may be performed prior to receiving the request 112, as described below with reference to Figure 2 Described in more detail.
[0071] In addition, despite Figure 1112, but different language models may be specifically trained to handle different prompts at different stages of the processing pipeline. For example, a more general (e.g., larger) language model may be used as an offline process (e.g., independent of receiving request 112) to generate a summary of online content, which may then be inserted into a prompt that is input to a more specialized and faster language model in an online process (e.g., in real time in response to receiving request 112). Additionally, AI system 160 may generate a set of candidate digital components as an offline process (e.g., before receiving request 112) and store the set of candidate digital components in a database. In this scenario, when AI system 160 receives request 112, AI system 160 may further evaluate and rank the stored candidate digital components based on additional information included in the request and other contextual data (e.g., time of day, date of week, weather conditions, etc.).
[0072] Figure 2 2 is a block diagram 200 illustrating the interaction between artificial intelligence system 160, language model 202, and client device 204. In some cases, language model 202 and client device 204 may be respectively Figure 1 The language model 170 is the same or similar to the client device 106. Figure 2 A single language model 202 is depicted in FIG, but the language model 202 may be a set of different language models that may be called upon for different tasks, with the different language models being specifically trained for the different tasks. For example, one language model within the set of language models may be specifically trained to perform a content summarization task, while another model may be specifically trained to generate highly factual outputs, for example, using the summary output of a specifically trained summary language model. In addition, the set of models may include a generalized language model that is larger in size and capable of generating a large and diverse set of data, but this generalized model may have a higher latency than the specialized models, which may make the generalized model less desirable for use in real-time operations depending on the time latency constraints required to generate content.
[0073] The artificial intelligence system 160 includes a data collection device 206, a summary device 208, a prompt device 210, and a post-processing device 212. The following description refers to these different devices as being independently implemented and each configured to perform a set of operations, but any of these devices may be combined to perform the operations discussed below.
[0074] The artificial intelligence system 160 communicates with a memory structure 214. The memory structure 214 may include one or more databases. As shown, the memory structure includes a collected data database 216, a clause database 218, and a digital component database 220. Each of these databases 216, 218, and 220 may be implemented in the same hardware memory device, separate hardware memory devices, and / or in a distributed cloud computing environment.
[0075] The data collection device 206 is implemented using at least one computing device (e.g., one or more processors) and may include one or more language models. The data collection device 206 is configured to collect information from online data sources. In some implementations, the collected information includes passages collected from a set of online resources. To obtain passages, the data collection device 206 may issue / submit a query to a search system, which responds to the query with information about topics and / or entities. In some implementations, the collected data obtained by the data collection device 206 may include search result fragments returned by the search system in response to submitting a query to the data collection device 206.
[0076] In some implementations, the data collection device 206 can be configured to rewrite / enhance the query and submit the rewritten query to the search system, for example, using the language model 202 or an internal language model. Performing this additional query rewriting and search process increases the diversity of the information collected, which is later summarized and ultimately used to generate clauses that will be used to create candidate digital components. Increasing the diversity of the information collected and ultimately used to generate clauses can increase the creative features of the candidate digital components by providing more output options for the language model while still complying with the specified constraints.
[0077] When an entity (e.g., a company) has an online presence (e.g., a website) that provides information about the entity, the query submitted by the data collection device 206 can be a site-constrained query, which causes the search system to reply to the site-constrained query with only information contained in a specified site (e.g., the company's website). Of course, multiple site-constrained queries can be issued for multiple different sites, or multiple different sites can be specified in a site-constrained query, which causes the search system to collect information related to the query from multiple different specified sites (e.g., social networking sites, web Q&A sites, entity review sites, etc.). The site constraint can be specified, for example, as a second-level domain or a specific page address, depending on where the information originates.
[0078] The data collection device 206 may store the collected data in a collected data database 216. For example, the data collection device 206 may index the collected data to a query used to collect the data and / or an entity represented by the collected data, so that the collected data may be retrieved from the collected data database 216 for use in additional operations performed by the data collection device 206 and / or any operations performed by the artificial intelligence system 160.
[0079] In some implementations, the collected data can be used to train or specifically train one or more language models. For example, a prompt related to the collected data can be submitted to a general language model, and the output of the general language model can be evaluated based on how closely the output of the general language model aligns with the collected data. More specifically, a quality metric can be calculated based on the difference (e.g., factual difference) between the output of the general language model and the collected data related to the prompt. The general language model can then be iteratively adjusted to reduce the difference between the output of the general language model and the collected data, which will result in a higher quality metric, thereby producing a specialized language model. This specialized language model can then be used to generate an output related to the collected data (e.g., a topic or content category for a subset of the collected data).
[0080] The summarization device 208 is implemented using at least one computing device (e.g., a device including one or more processors) and may include one or more language models. The summarization device 208 is configured to summarize information about a topic or entity (e.g., a person, place, thing, or concept). In some implementations, the summarization device 208 is configured to summarize data collected by the data collection device 206 and possibly stored in the collected data database 216. For example, the summarization device 208 can be configured to accept the collected data as input and output a specified length (e.g., 200 words or some other word count) summary of the content of the collected data.
[0081] The summary may be generated using the language model 202, which may be part of or in data communication with the summary device 208. In some implementations, the summary device 208 (or the prompt device 210 discussed below) may generate a summary prompt that is submitted to the language model 202 as an input prompt 222. The summary prompt may specify one or more of the following:
[0082] - the set of sources that should be used to generate the summary;
[0083] - details about the set of sources that the language model should consider when summarizing the content;
[0084] - a factual basis directive specifying that the language model should provide citations to sources used to generate the summary;
[0085] - summary constraints, which specify information that should not be included in the summary (e.g., information that is not directly supported by the set of sources); and
[0086] - formatting constraints that specify how the output of the language model should be formatted (e.g., as bullet points or in paragraph form, with or without an introduction summary, overall length (e.g., number of characters or separate clauses)); and
[0087] - Tone constraints that specify the tone of the output (e.g., creative, humorous, sad, serious, or from the perspective of a specified entity such as an artist, engineer, or storyteller).
[0088] A sample summary prompt might take the following form:
[0089] “Given a question and a list of sources, write a short summary that cites individual sources and summarizes all of them as comprehensively aspossible. Each source is independent and might repeat or contradicting content from other sources. The summary should be directly supported by the given sources and cited appropriately with a [$i] notation following a statement that is supported by $i. Thesummary may start with a general statement about the answer space. Thesummary shouldn't include any information that cannot be supported by thegiven sources.(Given the question and a list of sources, write a brief summary that cites the individual sources and summarizes all of them as comprehensively as possible. Each source is independent and may repeat or contradict content from other sources. The summary should be directly supported by the given sources and appropriately cited with the [$i] symbol after the statement supported by $i. If the statement is based on multiple sources, all of them should be listed in parentheses, for example, [$i, $j, $k]. The summary may begin with a general statement about the answer space. The summary should not include any information that is not supported by the given sources.)"
[0090] In this example, the symbol $i can serve as a placeholder for the name of the source. The bold "a list of sources" can be replaced with the name of the actual source to be considered, or as a reference to the location of the source in the set of sources to be considered when creating the summary. The set of sources can be the network address (e.g., universal resource indicator / locator-URI / URL) of an online data source (e.g., a second-level domain of a website, a specific address of a web page, or the network address of other data sources). In some implementations, the set of sources can include a collected data database 216 so that the summary can be generated using the data collected and stored by the data collection device 206. The summary device 208 uses summary prompts to generate a summary that summarizes the chapters collected from one or more network locations of the set of sources into a chapter summary. As noted above, the chapter summary can be formatted as a set of key points or in the form of paragraphs.
[0091] The summaries are generated by the language model 202, which, as noted above, may be part of or in data communication with the summarization device 208. In either case, the summarization device 208 inputs the summary prompt into the language model 202 as an input prompt 222. The language model 202 (e.g., LLM) processes the input prompt 222 and generates a natural language output ("NL output") 224 that summarizes the content of the set of sources according to the instructions / constraints specified in the summary prompt.
[0092] An example paragraph summary of a set of sources providing information about cars might take the following form:
[0093] The 2.5-liter 4-cylinder engine provides increased horsepower and torque distribution, allowing you to feel seamless acceleration. State-of-the-art all-wheel drive technology comes with six available driving modes, namely Economy, Standard, Tarmac, Gravel, Snow, Mud, so you can drive with confidence in almost any conditions. Active Blind Spot Assist can help drivers avoid collisions by detecting other vehicles in surrounding lanes, even when they are not visible to the human eye.
[0094] An example bullet point summary for the same episode of the source might take the following form:
[0095] - The 2.5L 4-cylinder engine provides increased horsepower and torque distribution, allowing you to feel seamless acceleration.
[0096] - State-of-the-art all-wheel drive technology with six available driving modes, Eco, Standard, Tarmac, Gravel, Snow, Mud, so you can drive with confidence in almost any conditions.
[0097] - Active Blind Spot Assist can help drivers avoid collisions by detecting other vehicles in surrounding lanes, even when they are not visible to the human eye.
[0098] The summary may be generated in response to receiving the query 226 from the client device 204 (e.g., in real-time or online mode), or in an offline mode (e.g., independent of the instance in which the query 226 is received from the client device 204). In the real-time / online mode, the query 226 received from the client device 204 may be passed to the summarization device 208 in parallel with other processing performed on the query (such as obtaining search results or otherwise generating a language model response to the query), so that the summarization device 208 can generate the summary while performing other query processing operations, thereby reducing the latency associated with providing the final response to the query 226 to the client device 204.
[0099] In offline mode, the summary device 208 can generate a summary of expected queries (e.g., high-volume queries), which can be stored in a memory structure for use when a query 226 is received from the client device 204. For example, the summary device 208 can identify a set of queries that are ranked highest because it is related to query frequency or another metric (such as time sensitivity of responses to queries). In this example, the summary device 208 can identify, for each query, a set of sources related to the query and / or stored collected data related to the query (e.g., collected using the query or similar queries), and perform operations similar to those discussed above to generate a set of summaries of the query (e.g., one or more summaries). This set of summaries can be stored in a memory structure 214 (e.g., with an index pointing to the query), and when the client device 204 submits a query 226, the summary device 208 (or another device in the AI system 160) can query the memory structure 214 to retrieve one or more summaries indexed to the query 226 to facilitate operations performed using the summaries, as discussed in more detail below. Generating summaries in offline mode may reduce latency associated with responding to queries 226 submitted by client devices 204 because the operations required to generate the summaries will not interfere with downstream operations discussed below that rely on the summaries (eg, prompt generation).
[0100] In some implementations, the summarization device 208 may perform one or more post-processing operations on the summaries output by the language model. The post-processing operations may evaluate the summaries based on one or more of their factuality, relevance, comprehensiveness, conciseness, tone, and clarity, as described in more detail below with reference to the post-processing device 212. Of course, in some implementations, the post-processing operations on the summaries may be performed by the post-processing device 212 itself.
[0101] The summary is provided to a prompting device 210, which is implemented using at least one computing device (e.g., a device including one or more processors) and may include one or more language models. The prompting device 210 is configured to generate a prompt, which includes a query 226 and a set of constraints.
[0102] For example, query 226 may be received from client device 204. Query 226 may be input through a search service, a chat interface, a game interface, a digital assistant interface, or another interface of a service provided online or provided by a native application installed at the client device. Query 226 may be as simple as a single word-gram, or may be a series of word-grams that constitute a multi-word-gram term. In this scenario, query 226 is received by AI system 160 and may be inserted into a prompt by prompt device 210. Additionally or alternatively, AI system 160 may use query 226 to search for or otherwise obtain information related to query 226. For example, AI system 160 may use query 226 to identify relevant information in stored collected data database 216, collect data related to query 226 from various online locations as described above with reference to data collection device 206, or otherwise use query 226 to generate or identify information that may provide additional context for the creation of a prompt (e.g., collect weather information related to the query, etc.).
[0103] The set of constraints may include a chapter summary for a specified source of online content, the chapter summary summarizing a chapter of information collected from the set of online resources (e.g., as described above with reference to the data collection device 206). For example, the prompt device 210 may insert one or more of the summaries generated by the summary device 208 into the prompt. In some implementations, the chapter summary inserted into the prompt operates as a contextual constraint that limits the content created by the language model 202 in response to the prompt containing the summary. For example, the summary may limit the content created by the language model to the topic specified by the summary, which is included in the prompt as a contextual constraint, as described in more detail below.
[0104] When constructing a prompt, the prompt device 210 can insert a specified entity name into the prompt. The entity name can be pre-specified or derived. In some implementations, the entity name is pre-specified based on the entity for which the content is being created (e.g., a content distributor). For example, assume that Example_Entity_1 ("ET1") has submitted a request to generate content to the AI system 160. In this example, the prompt device 210 can insert "ET1" into the prompt, which will be submitted to the language model 202 as an input prompt 222 to inform the language model 202 of the entity to be referenced in the NL output 224 of the language model 202. For example, Figure 1 Similar to the discussion of the example additional prompt 172, in this example, the prompt device 210 can insert the entity name "ET1" as an entity constraint of the input prompt 222 submitted to the language model 202. In this example, the entity constraint "ET1" acts as an indication to the language model 202 that the NL output 224 should reference "ET1". In some implementations, the prompt device 210 can insert a specific instruction into the prompt, the specific instruction being that the NL output 224 generated by the language model 202 must include the entity name inserted into the prompt.
[0105] In some implementations, entity names can be derived. For example, the prompt device 210 (or another device in the AI system 160) can evaluate various data to determine the appropriate entity specified in the entity constraint inserted into the prompt. The evaluated data may include, for example, the source of the data specified in the summary prompt, the summary itself, the collected data stored in the collected data database 216, or other sources of information. For illustration, assume that the source used to generate the summary mentions ET1 more frequently than any other entity. In this example, the prompt device 210 can determine that ET1 should be explicitly mentioned by the NL output 224 of the language model based on the fact that ET1 is mentioned more frequently than any other entity. In this example, the prompt device 210 derives the identity of the entity to be referenced in the entity constraint of the prompt by analyzing the source of information from which the summary is created, rather than being explicitly instructed to include a specific entity name in the prompt. This derived entity name can be inserted into the prompt, which will instruct / cause the language model 202 to reference the entity in the NL output 224 (e.g., specify the entity name in a textual manner).
[0106] The prompt device 210 can be configured to insert one or more basis constraints into the prompt. The basis constraints instruct / cause the language model to generate output from a verifiable source (e.g., a set of accessible sources). Including the basis constraints in the input prompt 222 can require that the content created by the language model 202 exists in a specified set of online resources or other data sources. To utilize the basis constraints, the prompt device 210 can insert instructions into the prompt, the instructions, i.e., the NL output 224 of the language model 202, only includes information from the summary (e.g., information from the summary with corresponding quotations) and / or actually / currently exists in the quotations referenced in the summary.
[0107] Including this basis constraint enables the language model 202 to construct an NL output 224 that can be verified within a specified data source (e.g., a summary, website, or other data source). This can produce a more factual and / or accurate NL output 224 that is less prone to "hallucinations," thereby improving the operation, accuracy, and / or precision of the language model 160 and the entire AI system 160. Language model hallucination refers to a situation in which a large language model (LLM) generates text that is not supported by the source data, is factually incorrect, or meaningless (e.g., stating that a dog has antlers). The nature of language models allows them to generate inaccurate, imprecise, or meaningless outputs in the absence of constraints. However, by using the basis constraints discussed herein, these types of erroneous or meaningless outputs can be avoided, thereby making the output of the language model more accurate, precise, and reliable. In addition, generating erroneous or meaningless data wastes computing resources, network bandwidth, mobile client battery power, and the like. Reducing the likelihood or occurrence of hallucinations (e.g., using grounds constraints) improves the operation of language models and systems that rely on the output of the language models, for example, by reducing the distribution of non-factual information, reducing the number of network calls to / from the language model to derive appropriate answers, and using less computing power to generate erroneous information. Reducing the likelihood / occurrence of model hallucinations also results in more efficient use of mobile device battery consumption, as the mobile device does not waste processing power or battery consumption processing and displaying hallucinations, which may also result in more queries from the user and more responses to be processed and displayed by the client device before deriving a factual response.
[0108] In some implementations, a specific network location (e.g., a second-level domain, such as example.com, or a full page address, such as example.com / example_page) or a set of network locations may be specified in accordance with the constraints. In these implementations, the prompting device 210 may insert the specific network location or set of network locations into the prompts provided to the language model as input prompts 222. Including these network locations in the prompts may instruct / cause the language model 202 to generate content that exists in one or more of the specified network locations, thereby increasing the likelihood that the NL output 224 generated by the language model is in fact accurate.
[0109] The prompt device 210 (or another component of the AI system 160) sends, conveys, communicates, or otherwise submits the constructed prompt to the language model 202. The language model 202 uses any of the summary, the query, and the specified constraints to generate an NL output 224. The NL output 224 can be a set of clauses formatted according to the formatting constraints specified in the input prompt 222. For example, if the input prompt 222 includes a formatting constraint specifying a "list of bullet points", the NL output 224 can be in the form of a bullet point list of clauses. The number of clauses included in the NL output 224 can also be specified by the constraints included in the prompt.
[0110] In some implementations, the number of clauses generated by language model 202 and included in NL output 224 may be higher than the number of clauses that will be used by AI system 160 to create each candidate digit component. For example, assume that AI system 160 will create candidate digit components that each include 3 clauses. In this example, AI system 160 may include in the prompt an instruction for language model 202 to generate at least 12 clauses.
[0111] By instructing the language model 202 to generate more clauses (e.g., sentences, key words, etc.) than are required to generate separate candidate digital components, the AI system 160 will be able to create multiple different digital components while only having to submit a single input prompt 222 and receive / process a single NL output 224. In this way, the system is more efficient than a system that requires multiple input prompts 222 and multiple single NL outputs 224 to create multiple candidate digital components. For example, by requesting 12 clauses in a single input prompt 222, the AI system 160 will receive 12 clauses in a single NL output 224, which can be used to create 220 different combinations of three clauses. Therefore, a single NL output 224 in this example can be used to create a minimum of 220 different candidate digital components, whereas if the NL output 224 only includes three clauses, the AI system 160 will only use one set of clauses to create the candidate digital component.
[0112] The AI system 160 may also use other objects to create additional candidate digital components. For example, the AI system 160 may use formatting options (e.g., font, font color, text emphasis, etc.) to create additional candidate digital components. The AI system 160 may also use multiple different links to create different candidate digital components. For example, the AI system 160 may combine the output of the language model 202 (e.g., a set of clauses) with a link to a specific sub-page within a secondary domain or a secondary domain to create a candidate digital component. When the AI system 160 has access to multiple links suitable for a given set of clauses received from the language model 202, the AI system 160 may generate multiple different candidate digital components, each candidate digital component including the same given set of clauses, but linking to different web pages.
[0113] For example, assume that AI system 160 identifies a link to a home page of a website (e.g., example.com) and a product information page (e.g., example.com / product_info), and that the clause of the digital component describes a product found at the product information page. In this example, AI system 160 may create one candidate digital component that links to the home page and another candidate digital component that links to the product information page, while using the same clause in each of the candidate digital components.
[0114] The clauses obtained from the language model 202 may be stored in the clause database 218 for further processing by the post-processing device 212 .
[0115] The post-processing device 212 of the AI system 160 is implemented using at least one computing device (e.g., a device including one or more processors) and may include one or more language models. The post-processing device 212 is configured (e.g., specially programmed with code) to perform one or more post-processing operations on the candidate digital component. In some implementations, the post-processing operation may occur after the digital component has been constructed (e.g., combining clauses, links, and / or other objects into candidate digital components). In some implementations, the post-processing operation may be performed before the construction of the candidate digital component is completed. For example, one or more of the post-processing operations may be performed on the clauses in the NL output 224 of the language model 202 before combining the clauses with the links to create a complete candidate digital component. As used throughout this specification, unless otherwise specified, performing a post-processing operation on a clause before combining into a complete candidate digital component is considered to be performing post-processing on the candidate digital component.
[0116] Performing post-processing operations includes evaluating one or more characteristics of the candidate digital component. The evaluated characteristics may include, for example, the factuality of the clause generated by the language model 202, the relevance of the clause to the query or summary included in the prompt, the completeness level of the clause, and the tone of the clause. The post-processing operations may be performed on clauses individually, on multiple clauses in combination, and / or on a complete candidate digital component that includes other objects (such as links to online resources).
[0117] In some implementations, the post-processing device 212 is configured to generate a basis score for each clause output by the language model 202. The basis score indicates the possibility that the clause is factual. For example, a clause existing in the current version of the specified online resource or data source can be regarded as more factual than a clause not existing in the current version of the specified online resource or data source. Since the language model uses summaries, queries, and constraints to generate new text content, many clauses are likely to be unable to be found verbatim in the specified online resource. Therefore, a similarity measure and / or another language model can be used to perform factual evaluation. For example, the clause can be compared with the original text of the specified online resource or data source to determine the semantic distance between the clause and the text of the specified online resource. Similarly, another language model can be used to determine the similarity level between the clause and the text of the online resource or data source. The post-processing device 212 can generate a basis score based on the similarity between the clause and the text of the specified online resource or data source, wherein a higher similarity level corresponds to a higher factual possibility and a higher basis score, and a lower similarity level corresponds to a lower factual level and a lower basis score.
[0118] The post-processing device 202 may be configured to filter the clauses output by the language model 202 based on the basis score. For example, if the basis score of the one or more clauses fails to meet the basis threshold (e.g., the minimum specified basis score), the post-processing device 202 may remove one or more clauses from consideration for inclusion in the digital component. The basis threshold may be specified by an administrator or designer of the AI system 160 and may be a basis score that demarcates between clauses classified as factual and non-factual. Using scores to demarcate between factual and non-factual clauses eliminates the subjectivity associated with human assessments of factuality and non-factuality, thereby making it an objective assessment. When a clause is removed from consideration for inclusion in the digital component, the AI system 160 may identify another available clause to consider including it or request another set of clauses from the language model 202. For example, if the clauses are initially ranked based on relevance (e.g., in the NL output 224), the next highest ranked clause may replace the removed clause when creating a candidate digital component. Similarly, if a clause is removed from the complete candidate digital component, another clause can be selected to replace the removed clause.
[0119] The post-processing device 212 can be configured to evaluate the relevance of each of the clauses to the query or summary included in the prompt. In some implementations, the relevance of the clause to the query or summary can be determined by embedding the clause, query and / or summary in a multidimensional semantic space and determining the cosine distance between the embeddings. The relevance of the clause to the query or summary can also be determined by inputting the clause, query and / or summary into a machine learning model (e.g., a neural network) that has been trained to determine the semantic relevance between sets of text. The post-processing device 212 can generate a relevance score (e.g., wherein a higher score indicates a higher relevance level) for each clause based on the analysis, and rank them based on the relevance of the clause.
[0120] The post-processing device 212 can be configured to evaluate the completeness level of each of the clauses. The completeness level of a clause or a set of clauses specifies how fully a clause or a set of clauses describes one or more topics. In some implementations, the completeness level specifies how fully a clause or a set of clauses describes topics in the secondary domain used to generate a summary. Clauses can be evaluated before creating candidate digital components, and ranked, and / or evaluated together as a group in candidate digital components. In some implementations, the evaluation specifies how fully the set of clauses in the candidate digital component describes topics in the secondary domain (or at a specific page) linked to by a given candidate digital component. The completeness level may be higher when the set of clauses in the candidate digital component (e.g., 3 clauses) more fully describes the topic in the secondary domain (e.g., provides more details), and may be lower when the set of clauses is less fully described. The post-processing device 212 can generate a completeness score for each set of clauses and / or each candidate digital component, and rank the set of clauses / candidate digital components by the completeness score.
[0121] The post-processing device 212 can be configured to evaluate the tone of each or a set of clauses in the clause. The tone of a clause is an indication of whether the clause characterizes an item in a positive or negative manner. In some implementations, the positivity level or the negativity level can be used to generate a tone score, for example, a positive tone clause has a higher tone score (e.g., a positive score) than a neutral and negative tone clause, and a negative tone clause has a lower tone score (e.g., a negative score) than a neutral and positive tone clause. For example, a zero score can be assigned to a neutral tone clause so that they do not have a positive or negative impact on the overall tone of the candidate digital component.
[0122] For example, the tone of a clause can be generated by submitting the clause to a language model and asking the language model whether the tone is positive, neutral, or negative. Additionally or alternatively, the clause can be input into a machine learning model that has been trained (e.g., using labeled data) to classify the clause as positive, neutral, or negative by tone. The classification of the clauses can be used to assign a tone score to each clause, and the overall tone of the candidate digital component can be determined by aggregating (e.g., summing) the tone scores of the individual clauses. The post-processing device 212 can rank the clauses / candidate digital components based on the tone scores.
[0123] The post-processing device 212 may be configured to rank the candidate digital components based on one or more of the post-processing operations. For example, the post-processing device 212 may rank each of the candidate digital components based on any one of the scores / evaluations discussed above or a combination of the scores / evaluations discussed above. For example, the post-processing device may sum or average a plurality of different scores to obtain an aggregate score for a clause, a set of clauses, or a candidate digital component. In some implementations, the score may be weighted based on the relative importance of each evaluation to obtain an aggregate score (e.g., a weighted average). Using the aggregate score, the post-processing device 212 may rank a clause, a set of clauses, or a candidate digital component, and may identify one or more of the highest-ranked candidate digital components as one or more output digital components ("output DC") 228, which are supplied to the client device by the AI system 160. In some implementations, the creation and supply of the output digital component 228 is performed after receiving the query 226. In some implementations, the output digital components 228 may be generated in an offline process (e.g., prior to receiving the query 226) and stored in the digital component database 220 until the query 226 is received. At this point, one or more of the output digital components 228 may be retrieved from the digital component database 220 and served to the client device 204.
[0124] Figure 3 is a flow chart of an example process 300 for creating and supplying digital components using artificial intelligence. The operations of process 300 may be performed, for example, by Figure 1 The process 300 may be performed by a service device 110 (e.g., including an AI system 160 and / or a language model 170) or another data processing device. The operations of process 300 may also be implemented as instructions stored on a computer-readable medium, which may be non-transitory. Execution of the instructions by one or more data processing devices causes the one or more data processing devices to perform the operations of process 300.
[0125] Chapters are collected from a collection of online resources (302). As discussed above, chapters can be collected in a variety of ways. For example, chapters can be collected using a site-constrained query that requires that chapters be collected from one or more network locations specified in site constraints. In a specific example, a site-constrained query can include a site constraint (e.g., a query parameter) that specifies that content searches using the query must be limited to locations within a second-level domain (e.g., example.com) specified in the site constraint or to a specific page within the second-level domain (e.g., an item detail page).
[0126] Chapters may be collected in an offline process (eg, independent of and / or prior to receiving a query from a client device) or in an online process performed while processing a query received from a client device. Additional details of chapter collection are provided above with reference to data collection device 206.
[0127] The collected passages are summarized into passage summaries (304). As discussed above, summarization of passages collected from one or more network locations can be performed by a language model (such as a large language model) trained to summarize multiple passages of text. In some implementations, passages are collected from a list of sources specified in a summary prompt generated by an artificial intelligence system and submitted to a large language model, as described above.
[0128] Summary hints submitted to a large language model can specify one or more of the following:
[0129] - the set of sources that should be used to generate the summary;
[0130] - details about the set of sources that the language model should consider when summarizing the content;
[0131] - a factual basis directive specifying that the language model should provide citations to sources used to generate the summary;
[0132] - summary constraints, which specify information that should not be included in the summary (e.g., information that is not directly supported by the set of sources); and
[0133] - formatting constraints that specify how the output of the language model should be formatted (e.g., as bullet points or in paragraph form, with or without an introduction summary, overall length (e.g., number of characters or separate clauses)); and
[0134] - Tone constraints that specify the tone of the output (e.g., creative, humorous, sad, serious, or from the perspective of a specified entity such as an artist, engineer, or storyteller).
[0135] Example summary prompts and additional details regarding the generation of summaries are provided above with reference to summary device 208 .
[0136] Generate / build a prompt (e.g., an additional prompt) that includes a query and a set of constraints that restrict clauses generated by the language model (306). A prompt that includes a query and a set of constraints is different from a summary prompt, and as described above, the set of constraints can include a summary generated using the summary prompt as a summary constraint. For example, generation of the prompt can include inserting at least a portion of the passage summary into the prompt as a context constraint that restricts content created by the language model to the topics specified in the context constraint.
[0137] Additionally, generation of the prompt may include inserting an entity name of an entity referenced by one or more network locations into the prompt as an entity constraint. The entity constraint specifies that content identifying the entity must be included in content created by the language model. In other words, the entity constraint instructs / causes the language model to include content identifying the entity in content (e.g., a clause) created by the language model.
[0138] The generation of the prompt may include inserting a basis constraint into the prompt, the basis constraint requiring that the content created by the language model exists in a specified set of online resources. For example, the insertion of the basis constraint into the prompt may be implemented by inserting a secondary domain (e.g., example.com) into the prompt, the secondary domain requiring that the content created by the language model exists in resources within the secondary domain. Of course, the network location of other data sources (e.g., databases, such as the collected data database 216, separate web pages, etc.) may also be inserted into the prompt as a basis constraint. The generation of the prompt is discussed in more detail above with respect to the prompt device 210.
[0139] A plurality of candidate digital components are generated / constructed using the clauses generated by the language model (308). In some implementations, the clauses are generated using a summary of a specified source of online content. The candidate digital component can be a single clause obtained from the language model or a combination of clauses obtained from the language model. The candidate digital component can also include other objects / items, such as links to online resources, scripts that implement various user interactions with the digital component (e.g., making a reservation, launching a game, launching an augmented reality environment, etc.). For example, one or more of the candidate digital components can be generated by combining the output of the language model (e.g., one or more clauses) with a link to a secondary domain (e.g., the homepage of example.com) and / or a link to a specific page within the secondary domain (e.g., an item information page for the item described by the clause).
[0140] In some implementations, each different candidate digital component includes a different combination of clauses received from the language model. For example, as discussed above, if the output of the language model (e.g., generated using the prompt from operation 306) includes 12 different clauses, and each of the digital components being generated by the AI system is formatted to include space for three different clauses, the AI system can use 3 different clauses in each of the candidate digital components to make 220 different candidate digital components (e.g., 12! / (3!(12-3)!)=220). Of course, adding other sets of objects / items to the digital components will further expand the possible number of combinations.
[0141] One or more post-processing operations are performed (310). In some implementations, the one or more post-processing operations include operations that evaluate one or more characteristics of each given candidate digital component in a plurality of different candidate digital components. As noted above, the candidate digital components can be single clauses output from the language model, combinations of clauses, and / or other objects that are combined with one or more of the clauses output from the language model. Thus, the post-processing operations can be performed on any of these candidate digital components (including individual clauses).
[0142] Performing one or more post-processing operations can be accomplished by evaluating the degree of factuality of the candidate digital component. In some implementations, the degree of factuality of the candidate digital component can be evaluated based on whether the information within the candidate digital component can be verified at one or more specified data sources.
[0143] For example, assume that the digital component is using a plurality of clauses generated by a language model to describe an object. In this example, information about the object would have been collected from a set of online resources as described above with reference to operation 302, summarized as described above with reference to operation 304, and then used by the language model to generate clauses output by the language model. The clauses output from the language model may be different from the passages collected in operation 302, for example, in order to present information from the passages in a more creative manner. Thus, the clauses may not be found verbatim in the set of online resources, but the clauses may still be analyzed to determine whether the information conveyed by the clauses is consistent with the information conveyed by the original passages, as described in more detail above with reference to post-processing device 212.
[0144] In some implementations, a basis score may be used to evaluate the degree of factuality of a candidate digital component (e.g., a single clause or a combination of clauses), as described above with reference to the post-processing device 212. For example, for each clause of the output of the language model, a basis score may be generated for the likelihood that the specified clause is factual based on the level of similarity / difference between the clause and the content of a specified online resource or data source.
[0145] Using the basis scores, one or more clauses can be filtered out (e.g., removed from consideration for supply in a candidate digital component). For example, one or more clauses having a basis score that fails to meet a basis threshold can be removed from consideration. The basis threshold is specified to delineate between clauses that are classified as factual and non-factual. Using a specified basis threshold (e.g., a minimum score) based on a semantic distance (e.g., a cosine distance) between a clause and reference content (e.g., at a specified online resource) eliminates the subjectivity of whether information is factual or non-factual, thereby producing an objective classification system. When one or more clauses are removed due to failure to meet the basis threshold, the one or more clauses can be removed and replaced with another clause in the output having a basis score that meets the basis threshold, or another clause can be evaluated for inclusion in the set of clauses considered for inclusion in the digital component.
[0146] As discussed above with reference to the post-processing device 212, the post-processing operation may include an evaluation of other characteristics of the candidate digital components. For example, each given candidate digital component in the plurality of candidate digital components may be evaluated with respect to relevance, completeness, and tone, among others. The evaluation of relevance may include evaluating the relevance of clauses in a given candidate digital component to one or more of the content of the prompted query, the prompted summary, a search result fragment generated using the prompted query, or a set of online resources from which the passage was collected (or another specified online data source).
[0147] The evaluation of the level of completeness specifies how comprehensively the clauses in a given candidate digital component describe one or more topics. As previously discussed, one or more topics may be those found in the secondary domain linked to or used to generate a summary by a given candidate digital component. The level of completeness may be higher when the set of clauses in the candidate digital component (e.g., 3 clauses) more fully describes the topic in the secondary domain (e.g., provides more details found in the secondary domain), and may be lower when the set of clauses is less fully describing the topic. For example, an artificial intelligence agent / machine learning system can compare the semantic space covered by the content of the secondary domain (e.g., in a multidimensional semantic space) with the semantic space covered by the set of clauses. The difference (e.g., mathematical difference or ratio) between the covered semantic spaces can be used to derive the completeness score of the set of clauses. The difference between the semantic spaces covered by different sets of content can be determined, for example, by embedding the text of the content (e.g., represented by a vector) and determining the distance between the embeddings (or the overlap level between the embeddings). Additionally or alternatively, different sets of content can be input into a neural network trained to determine semantic similarity.
[0148] Post-processing operations may include evaluating the tone of clauses of the candidate digital component to determine whether the clauses characterize the item in a positive or negative tone. In some implementations, the level of positivity or the level of negativity may be used to generate a tone score, for example, a positive tone clause has a higher tone score (e.g., a positive score) than a neutral and negative tone clause, while a negative tone clause has a lower tone score (e.g., a negative score) than a neutral and positive tone clause. For example, a zero score may be assigned to neutral tone clauses so that they do not have a positive or negative impact on the overall tone of the candidate digital component.
[0149] For example, the tone of a clause can be generated by submitting the clause to a language model and asking the language model whether the tone is positive, neutral, or negative. Additionally or alternatively, the clause can be input into a machine learning model that has been trained (e.g., using labeled data) to classify the clause as positive, neutral, or negative by tone. The classification of the clauses can be used to assign a tone score to each clause, and the overall tone of the candidate digital component can be determined by aggregating (e.g., summing) the tone scores of the individual clauses.
[0150] Each of the plurality of candidate digital components is ranked (312). In some implementations, the plurality of candidate digital components may be ranked based on the results of a post-processing operation. The candidate digital components may be ranked based on any of the scores / evaluations discussed above or a combination of the scores / evaluations discussed above. For example, the post-processing device may sum or average a plurality of different scores to obtain an aggregate score for a clause, a set of clauses, or a candidate digital component. In some implementations, the scores may be weighted based on the relative importance of each evaluation to obtain an aggregate score (e.g., a weighted average), which may be determined by a system administrator, a system architect, and / or a machine learning model that evaluates performance feedback of the candidate digital components. Using the aggregate score, the clause, a set of clauses, or a candidate digital component may be ranked (e.g., from highest score to lowest score).
[0151] At least one output digital component is supplied based on the ranking (314). The at least one output digital component may be selected, for example, from the highest ranked candidate digital components that may be classified as output digital components. More specifically, if one output digital component is to be supplied, the highest ranked output digital component may be supplied. If more than one output digital component is to be supplied, a set of multiple output digital components within the set of the highest ranked digital components may be supplied. Supplying the output digital component may include sending an instruction to cause the output digital component to be presented at the client device.
[0152] Figure 44 is a block diagram of an example computer system 400 that can be used to perform the operations described above. System 400 includes a processor 410, a memory 420, a storage device 430, and an input / output device 440. Each of components 410, 420, 430, and 440 can be interconnected, for example, using a system bus 450. Processor 410 is capable of processing instructions for execution within system 400. In one implementation, processor 410 is a single-threaded processor. In another implementation, processor 410 is a multi-threaded processor. Processor 410 is capable of processing instructions stored in memory 420 or on storage device 430.
[0153] The memory 420 stores information within the system 400. In one implementation, the memory 420 is a computer readable medium. In one implementation, the memory 420 is a volatile memory unit. In another implementation, the memory 420 is a non-volatile memory unit.
[0154] The storage device 430 can provide mass storage for the system 400. In one implementation, the storage device 430 is a computer-readable medium. In various implementations, the storage device 430 may include, for example, a hard disk device, an optical disk device, a storage device shared by multiple computing devices over a network (e.g., a cloud storage device), or some other mass storage device.
[0155] The input / output device 440 provides input / output operations for the system 400. In one implementation, the input / output device 440 may include one or more of a network interface device (e.g., an Ethernet card), a serial communication device (e.g., and an RS-232 port), and / or a wireless interface device (e.g., and an 802.11 card). In another implementation, the input / output device may include a driver device configured to receive input data and send output data to other devices (e.g., a keyboard, a printer, a display, and other peripheral devices 460). However, other implementations may also be used, such as a mobile computing device, a mobile communication device, a set-top TV client device, etc.
[0156] Although already Figure 4 An example processing system is described in the specification, but the subject matter and functional operations described in this specification may be implemented in other types of digital electronic circuit systems (circuitry) or in computer software, firmware or hardware (including the structures disclosed in this specification and their structural equivalents) or in a combination of one or more of them.
[0157] An electronic document (referred to simply as a document for brevity) does not necessarily correspond to a file. A document may be stored in a portion of a file that holds other documents, in a single file dedicated to the document in question, or in multiple coordinated files.
[0158] For situations in which the systems discussed herein collect and / or use personal information about a user, the user may be provided with an opportunity to enable / disable or control programs or features that may collect and / or use personal information (e.g., information about the user's social network, social actions or activities, the user's preferences, or the user's current location). Additionally, certain data may be processed in one or more ways before it is stored or used so that personally identifiable information associated with the user is removed. For example, the user's identity may be anonymized so that personally identifiable information about the user cannot be determined, or the user's geographic location may be generalized (such as to a city, zip code, or state level) in the case where location information is available so that the user's specific location cannot be determined.
[0159] The embodiments of the subject matter and operations described in this specification may be implemented in a digital electronic circuit system or in computer software, firmware or hardware (including the structures disclosed in this specification and their structural equivalents) or in a combination of one or more of them. The embodiments of the subject matter described in this specification may be implemented as one or more computer programs, that is, one or more modules of computer program instructions, which are encoded on a computer storage medium for execution by a data processing device or for controlling the operation of a data processing device. Alternatively or in addition, program instructions may be encoded on an artificially generated propagation signal (e.g., a machine-generated electrical, optical or electromagnetic signal), which is generated to encode information for transmission to a suitable receiver device for execution by a data processing device. The computer storage medium may be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them, or may be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. In addition, although a computer storage medium is not a propagation signal, a computer storage medium may be a source or destination of computer program instructions encoded in an artificially generated propagation signal. The computer storage media may also be, or be included in, one or more separate physical components or media (eg, multiple CDs, disks, or other storage devices).
[0160] The operations described in this specification may be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0161] The term "data processing equipment" encompasses all types of equipment, devices and machines for processing data, for example, including a programmable processor, a computer, a system on a chip or a plurality of the foregoing or a combination thereof. The equipment may include a dedicated logic circuit system, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, the equipment may also include code for creating an execution environment for the computer program involved, for example, code constituting a processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine or a combination of one or more of them. The equipment and the execution environment may implement various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures.
[0162] This document relates to service devices. As used herein, a service device is one or more data processing devices that perform operations to facilitate the distribution of content over a network. A service device is depicted as a single block in a block diagram. However, while a service device can be a single device or a single set of devices, the present disclosure contemplates that a service device can also be a group of devices, or even multiple different systems that communicate to provide various content to client devices. For example, a service device can encompass one or more of a search system, a video streaming service, an audio streaming service, an email service, a navigation service, an advertising service, a gaming service, or any other service.
[0163] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network.
[0164] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by, and the device can also be implemented as, a special purpose logic circuit system, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0165] For example, processors suitable for the execution of computer programs include both general-purpose microprocessors and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic element of a computer is a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more large-capacity storage devices for storing data, such as a disk, a magneto-optical disk, or an optical disk, or be operatively coupled to receive data from the one or more large-capacity storage devices or to transfer data or both to the one or more large-capacity storage devices. However, the computer does not need to have such a device. In addition, the computer can be embedded in another device (e.g., a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name a few). Devices suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, by way of example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0166] To provide interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including sound, voice, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on a user's client device in response to a request received from the web browser.
[0167] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer with a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification), or any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include local area networks ("LANs") and wide area networks ("WANs"), interconnected networks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0168] A computing system may include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server sends data (e.g., an HTML page) to a client device (e.g., for the purpose of displaying data to a user interacting with the client device and receiving user input from the user). Data generated at the client device (e.g., the result of the user interaction) can be received from the client device at the server.
[0169] Although this specification contains many specific implementation details, these details should not be interpreted as limitations on the scope of any invention or the scope of what may be claimed, but rather as descriptions of features specific to a particular embodiment of a particular invention. Certain features described in this specification in the context of a separate embodiment may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable sub-combination. In addition, although features may be described above as working in certain combinations and even initially claimed as such, in some cases, one or more features from the claimed combination may be deleted from the combination, and the claimed combination may involve a sub-combination or a change in the sub-combination.
[0170] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be construed as requiring that such operations be performed in the particular order shown or in a sequential order, or that all of the operations shown be performed, in order to achieve the desired result. In certain contexts, multitasking and parallel processing may be advantageous. In addition, the separation of various system components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products.
[0171] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions set forth in the claims can be performed in a different order and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some implementations, multitasking and parallel processing may be advantageous.
Claims
1. A method comprising: generating, by an artificial intelligence system, a prompt comprising a query and a set of constraints limiting clauses generated by a language model, wherein the set of constraints comprises a summary of a specified source of online content; generating, by the artificial intelligence system, a plurality of candidate digital components using clauses generated by the language model using the summary of the specified source of online content; performing, by the artificial intelligence system, one or more post-processing operations to evaluate one or more characteristics of the plurality of candidate digital components; ranking, by the artificial intelligence system, each of the plurality of candidate digital components based on the post-processing operation; as well as At least one output digital component in the set of highest ranked candidate digital components is supplied by the artificial intelligence system.
2. The method of claim 1, further comprising: collecting passages from a collection of online resources using a site-constrained query that requires collecting the passages from one or more network locations specified in a site constraint; as well as The chapters collected from the one or more network locations are summarized into a chapter summary.
3. The method of claim 2, wherein generating the prompt comprises inserting at least a portion of the passage summary into the prompt as a context constraint, the context constraint limiting content created by the language model to topics specified in the context constraint.
4. The method of claim 3, wherein generating the hint comprises inserting an entity name of an entity referenced by the one or more network locations into the hint as an entity constraint, the entity constraint specifying that content identifying the entity must be included in content created by the language model.
5. The method of claim 4, wherein generating the hint comprises inserting a basis constraint into the hint, the basis constraint requiring that content created by the language model be present in a specified set of online resources.
6. The method of claim 5, wherein inserting a dependency constraint into the hint comprises inserting a secondary domain into the hint, the secondary domain requiring that content created by the language model be present in a resource within the secondary domain.
7. The method of claim 6, wherein generating the plurality of candidate digital components comprises combining an output of the language model with a link to the secondary domain.
8. The method of claim 1 , wherein performing one or more post-processing operations to evaluate one or more characteristics of the plurality of candidate digital components comprises: For each clause of the output of the language model, generating a basis score that specifies a likelihood that the clause is factual; filtering the output of the language model by removing one or more clauses having a basis score that fails to satisfy a basis threshold, the basis threshold demarcating between clauses classified as factual and non-factual; as well as One or more removed clauses are replaced with another clause of the output having a basis score that satisfies the basis threshold.
9. The method of claim 1 , wherein performing one or more post-processing operations to evaluate one or more characteristics of the plurality of candidate digital components comprises: For each given candidate numeric component in the plurality of candidate numeric components: evaluating the relevance of the clause in the given candidate digital component to the query of the prompt; evaluating a completeness level of the clause, the completeness level specifying how fully the clause in the given candidate digital component describes a topic in a secondary domain to which the given candidate digital component is linked; as well as The mood of the clause of the candidate numeric component is evaluated to determine whether the clause characterizes an item as positive or negative.
10. One or more non-transitory computer-readable media storing instructions that, when executed by an artificial intelligence system, cause the artificial intelligence system to perform operations comprising: generating a prompt comprising a query and a set of constraints restricting clauses generated by a language model, wherein the set of constraints comprises a summary of a specified source of online content; generating a plurality of candidate digital components using clauses generated by the language model using the summary of the specified source of online content; performing one or more post-processing operations that evaluate one or more characteristics of the plurality of candidate digital components; ranking each of the plurality of candidate digital components based on the post-processing operation; as well as At least one output digital component in the set of highest ranked candidate digital components is supplied.
11. The one or more non-transitory computer-readable media of claim 10, wherein the instructions cause the artificial intelligence system to perform operations further comprising: collecting passages from a collection of online resources using a site-constrained query that requires collecting the passages from one or more network locations specified in a site constraint; as well as The chapters collected from the one or more network locations are summarized into a chapter summary.
12. One or more non-transitory computer-readable media as described in claim 11, wherein generating the prompt includes inserting at least a portion of the passage summary into the prompt as a context constraint, the context constraint limiting content created by the language model to topics specified in the context constraint.
13. One or more non-transitory computer-readable media as described in claim 12, wherein generating the prompt includes inserting an entity name of an entity referenced by the one or more network locations into the prompt as an entity constraint, the entity constraint specifying that content identifying the entity must be included in the content created by the language model.
14. The one or more non-transitory computer-readable media of claim 13, wherein generating the hint comprises inserting a basis constraint into the hint, the basis constraint requiring that content created by the language model be present in a specified set of online resources.
15. One or more non-transitory computer-readable media as described in claim 14, wherein inserting a basis constraint into the prompt includes inserting a secondary domain into the prompt, the secondary domain requiring that content created by the language model exist in a resource within the secondary domain.
16. The one or more non-transitory computer-readable media of claim 15, wherein generating the plurality of candidate digital components comprises combining an output of the language model with a link to the secondary domain.
17. The one or more non-transitory computer-readable media of claim 10, wherein performing one or more post-processing operations to evaluate one or more characteristics of the plurality of candidate digital components comprises: For each clause of the output of the language model, generating a basis score that specifies a likelihood that the clause is factual; filtering the output of the language model by removing one or more clauses having a basis score that fails to satisfy a basis threshold, the basis threshold demarcating between clauses classified as factual and non-factual; as well as One or more removed clauses are replaced with another clause of the output having a basis score that satisfies the basis threshold.
18. The one or more non-transitory computer-readable media of claim 10, wherein performing one or more post-processing operations to evaluate one or more characteristics of the plurality of candidate digital components comprises: For each given candidate numeric component in the plurality of candidate numeric components: evaluating the relevance of the clause in the given candidate digital component to the query of the prompt; evaluating a completeness level of the clause, the completeness level specifying how fully the clause in the given candidate digital component describes a topic in a secondary domain to which the given candidate digital component is linked; as well as The mood of the clause of the candidate numeric component is evaluated to determine whether the clause characterizes an item as positive or negative.
19. An artificial intelligence system comprising: one or more memory devices; as well as One or more computing devices configured to execute code comprising a set of instructions, wherein execution of the set of instructions causes the one or more computing devices to perform operations comprising: generating a prompt comprising a query and a set of constraints restricting clauses generated by a language model, wherein the set of constraints comprises a summary of a specified source of online content; generating a plurality of candidate digital components using clauses generated by the language model using the summary of the specified source of online content; performing one or more post-processing operations that evaluate one or more characteristics of the plurality of candidate digital components; ranking each of the plurality of candidate digital components based on the post-processing operation; and At least one output digital component in the set of highest ranked candidate digital components is supplied.
20. The system of claim 19, wherein the instructions cause the one or more computing devices to perform operations further comprising: collecting passages from a collection of online resources using a site-constrained query that requires collecting the passages from one or more network locations specified in a site constraint; as well as The chapters collected from the one or more network locations are summarized into a chapter summary.
21. The system of claim 20, wherein generating the prompt comprises inserting at least a portion of the passage summary into the prompt as a context constraint, the context constraint limiting content created by the language model to topics specified in the context constraint.
22. The system of claim 21, wherein generating the prompt comprises inserting an entity name of an entity referenced by the one or more network locations into the prompt as an entity constraint, the entity constraint specifying that content identifying the entity must be included in content created by the language model.
23. The system of claim 22, wherein generating the hint comprises inserting a basis constraint into the hint, the basis constraint requiring that content created by the language model be present in a specified set of online resources.
24. The system of claim 23, wherein inserting a dependency constraint into the hint comprises inserting a secondary domain into the hint, the secondary domain requiring that content created by the language model be present in a resource within the secondary domain.
25. The system of claim 24, wherein generating the plurality of candidate digital components comprises combining an output of the language model with a link to the secondary domain.
26. The system of claim 19, wherein performing one or more post-processing operations to evaluate one or more characteristics of the plurality of candidate digital components comprises: For each clause of the output of the language model, generating a basis score that specifies a likelihood that the clause is factual; filtering the output of the language model by removing one or more clauses having a basis score that fails to satisfy a basis threshold, the basis threshold demarcating between clauses classified as factual and non-factual; as well as One or more removed clauses are replaced with another clause of the output having a basis score that satisfies the basis threshold.
27. The system of claim 19, wherein performing one or more post-processing operations to evaluate one or more characteristics of the plurality of candidate digital components comprises: For each given candidate numeric component in the plurality of candidate numeric components: evaluating the relevance of the clause in the given candidate digital component to the query of the prompt; evaluating a completeness level of the clause, the completeness level specifying how fully the clause in the given candidate digital component describes a topic in a secondary domain to which the given candidate digital component is linked; as well as The mood of the clause of the candidate numeric component is evaluated to determine whether the clause characterizes an item as positive or negative.