Revision and attribution of output of text generation model
By searching and comparing the content generated by the text generation model with relevant evidence sources, more consistent text content is generated and attributing it, the problem of generating incorrect content in the existing technology is solved, and effective revision and reliability improvement of text content is achieved.
Patent Information
- Application Number
- CN202380072615.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-29
- Filing Date
- 2023-09-29
- Publication Date
- 2025-05-27
AI Technical Summary
Existing text generation models may produce incorrect or unsupported content when generating content, resulting in the output being unsuitable for certain tasks, requiring a way to verify and improve the reliability of generated content.
The text content generated by the first machine learning language model is obtained through the computing system, the evidence source related to the text content is retrieved, the second text content is generated based on the comparison between the text content and the evidence source, and the evidence source and the generated text content are attributable.
Effective revisions of machine-generated text are realized, making it consistent with external evidence, improving the reliability and applicability of generated content, while retaining the style and structure of the text.
Smart Images

Figure CN120051770A_ABST
Abstract
Description
[0001] Priority claim
[0002] This application is based on and claims priority to U.S. Provisional Application No. 63 / 411,428, filed on September 29, 2022, which is incorporated herein by reference. Technical Field
[0003] The present disclosure relates generally to machine learning. More specifically, the present disclosure relates to a machine learning method for providing revisions and attributions (e.g., citations) for text (e.g., machine-generated text from an existing large language model). Background Art
[0004] Text generation models (TGMs) are now the backbone of many AI systems. For example, large language models can achieve multi-step reasoning, planning, rich interaction with the outside world, and open-domain question answering comparable to traditional information retrieval systems.
[0005] Despite these advances, state-of-the-art TGMs may still generate incorrect or unsupported content in some cases. Due to this problem, the output of TGMs may not be suitable for certain tasks. Therefore, methods for verifying the reliability of generated content are needed. Summary of the invention
[0006] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned through practice of the embodiments.
[0007] An example aspect of the present disclosure provides an example computer-implemented method for improving the attribution of model-generated content. The method may include obtaining, by a computing system including one or more computing devices, first text content generated by a first machine-learned language model. The method may also include retrieving, by the computing system, one or more sources of evidence related to the first text content generated by the first machine-learned language model. The method may also include generating, by the computing system, second text content as an output of the second machine-learned language model based on a comparison between the first text content and the one or more sources of evidence. And the method may also include associating, by the computing system, one or more portions of the one or more sources of evidence as attribution of the first or second text content.
[0008] In some examples, retrieving, by the computing system, the one or more sources of evidence may include: generating, by the computing system, one or more queries based on the first textual content; and retrieving, by the computing system, the one or more sources of evidence based on the one or more queries.
[0009] In some instances, retrieving, by the computing system, one or more sources of evidence based on the one or more queries may include a web search.
[0010] In some examples, generating, by the computing system, one or more queries based on the first textual content may include prompting a third machine-learned language model.
[0011] In some examples, generating, by the computing system, one or more queries based on the first textual content may include: identifying one or more claims associated with the first textual content; and generating a query associated with at least one of the one or more claims.
[0012] In some instances, the example computer-implemented method may include prompting a third machine-learned language model to identify one or more claims associated with the first textual content.
[0013] In some instances, the example computer-implemented method may include prompting a third machine-learned language model to generate a query associated with at least one or more claims associated with the first textual content.
[0014] In some examples, prompting the third machine-learned language model can include few-shot prompting.
[0015] In some examples, prompting the third machine-learned language model can include a thought chaining prompt.
[0016] In some instances, an example computer-implemented method may include: extracting, by a computing system, one or more evidence excerpts from one or more evidence sources retrieved by the computing system; and comparing, by the computing system, the one or more evidence excerpts extracted by the computing system with the first text content.
[0017] In some instances, extracting, by a computing system, one or more evidence excerpts from one or more evidence sources retrieved by the computing system may include: determining one or more candidate evidence excerpts; for each candidate evidence excerpt, generating a corresponding score based on a query used to retrieve the one or more evidence sources; and selecting a subset of the candidate evidence excerpts based on the corresponding scores.
[0018] In some instances, the corresponding score may be a relevance score generated by a machine-learned query-document relevance model.
[0019] In some instances, one or more sources of evidence are text documents, and determining one or more candidate evidence excerpts may include running a sliding window across the respective text documents.
[0020] In some examples, selecting the subset of candidate evidence excerpts based on respective scores may include maximizing coverage of one or more claims.
[0021] In some instances, comparing, by the computing system, the one or more evidence excerpts extracted by the computing system with the first textual content may include determining one or more first levels of consistency between the one or more evidence excerpts and the first textual content regarding one or more claims associated with the first textual content.
[0022] In some examples, determining one or more first levels of consistency between the one or more evidential excerpts and the first textual content regarding one or more claims associated with the first textual content may include utilizing a machine-learned consistency model to evaluate the first textual content and the evidential excerpts.
[0023] In some examples, utilizing the machine-learned consistency model to evaluate the first text content and the evidence excerpt can include prompting the machine-learned consistency model with a thought chaining prompt.
[0024] In some examples, utilizing the machine-learned consistency model to evaluate the first textual content and the evidential excerpt may include prompting the machine-learned consistency model with a few-shot prompt.
[0025] In some instances, generating, by a computing system, second textual content as an output of a second machine-learned language model may include: identifying a corresponding claim among one or more claims associated with the first textual content, wherein a first level of consistency with respect to the corresponding claim indicates a lack of complete consistency; and generating the second textual content such that a second level of consistency with respect to the corresponding claim between the second textual content and one or more evidence excerpts indicates a higher degree of consistency than the first level of consistency with respect to the corresponding claim.
[0026] In some examples, generating, by the computing system, the second textual content as an output of the second machine-learned language model can include prompting the second machine-learned language model using a thought chaining prompt.
[0027] In some examples, generating, by the computing system, the second textual content as an output of the second machine-learned language model can include prompting the second machine-learned language model using a few-shot prompt.
[0028] In some instances, the example computer-implemented method may also include: determining an edit distance based on one or more differences between the first text content and the second text content; and determining whether to output the second text content based on the edit distance.
[0029] In some instances, the example computer-implemented method may also include: determining a consistency score between the second text content and one or more evidence excerpts extracted from one or more evidence sources; determining a retention score indicating similarity between the first text content and the second text content; and determining whether to output the second text content based on a combination of the consistency score and the retention score.
[0030] In some instances, one or more sources of evidence may be text documents.
[0031] Other example aspects of the present disclosure provide an example computing system. The example computing system may include one or more processors and one or more non-transitory computer-readable media. The one or more non-transitory computer-readable media may store instructions that are executable by one or more processors to cause the computing system to perform one or more operations. The operation may include obtaining a first text content generated by a first machine-learned language model. The operation may also include retrieving one or more sources of evidence related to the first text content generated by the first machine-learned language model. The operation may also include generating a second text content as an output of a second machine-learned language model based on a comparison between the first text content and the one or more sources of evidence. And the operation may also include associating one or more portions of the one or more sources of evidence as an attribution of the first or second text content.
[0032] Other example aspects of the present disclosure provide one or more example computer-readable media. The example computer-readable medium may store instructions that can be executed by a computing system to cause the computing system to perform one or more operations. The operation may include obtaining a first text content generated by a first machine-learned language model. The operation may also include retrieving one or more sources of evidence related to the first text content generated by the first machine-learned language model. The operation may also include generating a second text content as an output of a second machine-learned language model based on a comparison between the first text content and the one or more sources of evidence. And the operation may also include associating one or more portions of the one or more sources of evidence as an attribution of the first or second text content.
[0033] Other aspects of the disclosure relate to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.
[0034] These and other features, aspects and advantages of various embodiments of the present disclosure will be better understood with reference to the following description and appended claims.The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] A detailed discussion of the embodiments for those skilled in the art is set forth in this specification with reference to the accompanying drawings, in which:
[0036] Figure 1A A block diagram depicts an example computing system 100 that performs attribution on automatically generated text according to an example embodiment of the present disclosure.
[0037] Figure 1B A block diagram depicting an example computing device that performs attribution on automatically generated text according to an example embodiment of the present disclosure.
[0038] Figure 1C A block diagram depicting an example computing device that performs attribution on automatically generated text according to an example embodiment of the present disclosure.
[0039] Figure 2 A block diagram depicting an example system and workflow for editing and attributing machine-generated text according to an example embodiment of the present disclosure.
[0040] Figure 3 A block diagram depicting an example method for editing and / or attributing machine-generated text according to an example embodiment of the present disclosure.
[0041] Figure 4 A block diagram depicting an example method for retrieving evidential excerpts related to machine-generated text according to an example embodiment of the present disclosure.
[0042] Figure 5 A block diagram depicting an example method in which machine-generated text is compared to one or more retrieved evidence excerpts according to an example embodiment of the present disclosure.
[0043] Figure 6 A block diagram depicting an example method for editing machine-generated text according to an example embodiment of the present disclosure.
[0044] Figure 7 A block diagram depicting an example experiment and corresponding example output according to an example embodiment of the present disclosure, in which machine-generated text is edited and attributed to an evidential excerpt.
[0045] Figure 8 A block diagram depicting an example experiment and corresponding example output according to an example embodiment of the present disclosure, in which machine-generated text is edited.
[0046] Fig. 9 Depicts example few-shot prompts for prompting an example machine learning model in an example experiment in accordance with the present disclosure.
[0047] Fig.10is a flow chart illustrating an example method for training a model for machine learning according to an example implementation of aspects of the present disclosure.
[0048] Fig.11 is a block diagram of an example process flow for processing inputs to generate outputs using a machine learning model according to an example implementation of aspects of the present disclosure.
[0049] Fig.12 is a block diagram of an example sequence processing model according to an example implementation of aspects of the present disclosure.
[0050] Fig.13 is a block diagram of an example technique for populating an example input sequence for processing by a sequence processing model in accordance with an example implementation of aspects of the present disclosure.
[0051] Fig.14 is a block diagram of an example model development platform according to an example implementation of aspects of the present disclosure.
[0052] Fig.15 is a block diagram of an example training workflow for training a model for machine learning according to an example implementation of aspects of the present disclosure.
[0053] Fig.16 is a block diagram of an inference system for operating one or more machine learning models to perform inference according to an example implementation of aspects of the present disclosure.
[0054] Reference numerals repeated across multiple drawings are intended to identify like features in the various implementations. DETAILED DESCRIPTION
[0055] Overview
[0056] Existing language models (LMs) can perform some tasks such as question answering, reasoning, and dialogue well. However, they sometimes generate unsupported or misleading content. Therefore, in one example aspect of the present disclosure, a mechanism for attributing the content of a LM to external evidence is provided. In another example aspect of the present disclosure, a mechanism for editing the content of a LM to make it more consistent with external evidence is provided.
[0057] In some example aspects of the present disclosure, systems and methods are provided that, given an input text generated by a LM, can retrieve external evidence related to the text, and can then revise the text to make it consistent with the evidence while preserving qualities such as style or structure, so that the revised text can be used seamlessly in place of the original text.
[0058] In some exemplary aspects of the present disclosure, an existing language model may generate text content that is factually unreliable, and the system and method of the present disclosure may edit the content to detect and correct errors in the generated text content. The research phase may use a machine-learned query generator to ask questions about different aspects of the text, such as factual claims made by the text. Next, these queries may be used to retrieve evidence from a corpus of evidence sources, such as by performing a web search using a query to find relevant text documents. Next, a machine-learned relevance model may be used to extract a set of evidence excerpts from the retrieved evidence sources based on a machine-learned determination of which parts of the retrieved sources are most relevant to the generated query. Next, a machine-learned consistency model may compare the evidence excerpts with the generated text content, and any consistency or inconsistency may be detected. Therefore, it may be determined whether the factual claims in the generated text content are supported by the extracted excerpts; whether they contradict the extracted excerpts; or whether they are mentioned by the extracted excerpts. Next, the machine-learned editing model may correct any text content that contradicts the extracted excerpts, while retaining the content and structure of the generated text content in other aspects.
[0059] Separately or additionally, the systems and methods of the present disclosure may associate the extracted evidence excerpts with relevant portions of the generated text content. The systems and methods of the present disclosure may then output information about the associated evidence excerpts to a user or another computing system or process. In some instances, the input text content may be output unedited along with a separate attribution report that explains which factual claims are supported by the evidence; contradict the evidence; or are not mentioned by the retrieved evidence. In other example instances, the generated content may be edited to ensure factual accuracy, and the edited content may be accompanied by a separate attribution report. In other instances, attribution may be included as part of the edited content, for example, as an in-text citation. Although the preceding sentences are discussed with factual claims as an example, the systems and methods of the present disclosure may also be used for viewpoints and other aspects of the generated content (e.g., attribution reports that identify sources that are inconsistent with the expressed viewpoints, or identify general background information about the topic being discussed).
[0060] In some additional aspects of the present disclosure, systems and methods are provided for measuring the quality of edited text content and retrieved evidence excerpts. Automated systems and methods are provided for measuring whether edited text content is attributable to retrieved evidence sources. In some example experiments, it is shown that the provided methods are highly correlated with human judgment about attribution. Automated systems and methods are also provided for measuring how much of the original text content is retained within the edited text content. In some example experiments using these measurements, it is shown that the systems and methods of the present disclosure enable the editing of text content to be better consistent with evidence while retaining the original content better than previous methods.
[0061] The present disclosure provides some technical advantages and benefits that will be apparent to those skilled in the art. For example, the systems and methods provided in the present disclosure significantly improve attribution while otherwise preserving the original input to a greater extent than previous work.
[0062] In addition, the systems and methods of the present disclosure can be performed using only a small number of training examples, which can save the computational costs associated with training machine learning models. Therefore, the systems and methods of the present disclosure can perform computing tasks using less power and computing time than previous systems and methods.
[0063] Furthermore, the systems and methods of the present disclosure can be performed on text generated by an already existing language model, which can take advantage of the benefits of large language models that are expensive to train (e.g., trillion-parameter pre-trained models), but without the power and computational costs associated with retraining such large models. Thus, the systems and methods of the present disclosure can perform tasks that may not be accomplished without the help of large (e.g., trillion-parameter) models and that cannot be performed by existing large language models alone, using less power and other training costs than previous work.
[0064] In addition, the systems and methods of the present disclosure can provide revised text that can be consistent with external evidence while retaining qualities such as style or structure, enabling the revised text to be used seamlessly in place of the original text. This avoids any need to generate additional text after editing is performed, thereby saving the power and computing costs associated with generating additional text after editing.
[0065] In addition, the systems and methods of the present disclosure can generalize to many tasks better than previous work. Therefore, the systems and methods of the present disclosure enable the reuse of the systems and methods across multiple tasks, which can save power, computation time, and memory space associated with training, storing, and operating separate machine learning models for each task.
[0066] Furthermore, the systems and methods of the present disclosure can be implemented using already existing document retrieval tools, such as standard web search engines. Thus, the systems and methods of the present disclosure can take advantage of the benefits of powerful web search tools without using the power, computing time, and memory space that may be associated with training, storing, and operating task-specific evidence retrieval tools.
[0067] In addition, the systems and methods of the present disclosure can enable the use of smaller language models to achieve similar performance levels as larger models, thereby saving power, computational costs, and memory space associated with larger models. For example, it has been shown that larger language models (e.g., trillion parameter models) "hallucinate" less frequently than smaller models (e.g., multi-billion parameter models), but training larger models requires more power, more memory, and more processing time than smaller models. The systems and methods of the present disclosure can achieve the increased accuracy associated with larger models without the increased power, memory, and processing time associated with training larger models.
[0068] Referring now to the accompanying drawings, example embodiments of the present disclosure will be discussed in greater detail.
[0069] Example Apparatus and Systems
[0070] Figure 1A A block diagram depicts an example computing system 100 that performs attribution on automatically generated text according to an example embodiment of the present disclosure. System 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 that are communicatively coupled via a network 180.
[0071] The user computing device 102 may be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop computer), a mobile computing device (e.g., a smartphone or tablet computer), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0072] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or a plurality of processors operatively connected. The memory 114 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 may store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.
[0073] In some implementations, the user computing device 102 may store or include one or more models 120. For example, the text attribution model 120 may be or may otherwise include various machine learning models, such as a neural network (e.g., a deep neural network) or other types of machine learning models, including nonlinear models and / or linear models. The neural network may include a feedforward neural network, a recurrent neural network (e.g., a long short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks. Some example machine learning models may utilize attention mechanisms, such as self-attention. For example, some example machine learning models may include a multi-head self-attention model (e.g., a transformer model).
[0074] In some implementations, one or more text attribution models 120 may be received from the server computing system 130 via the network 180, stored in the user computing device memory 114, and then used or otherwise implemented by the one or more processors 112. In some implementations, the user computing device 102 may implement multiple parallel instances of a single text attribution model 120 (e.g., to perform parallel attribution on automatically generated text across multiple attribution instances of the automatically generated text).
[0075] Additionally or alternatively, one or more text attribution models 140 may be included in or otherwise stored and implemented by a server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the text attribution models 140 may be implemented by the server computing system 140 as part of a web service (e.g., an automatically generated attribution of text service). Thus, one or more models 120 may be stored and implemented at the user computing device 102, and / or one or more models 140 may be stored and implemented at the server computing system 130.
[0076] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component may be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other devices that a user may use to provide user input.
[0077] The server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or a plurality of processors operatively connected. The memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 may store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.
[0078] In some implementations, server computing system 130 includes, or is otherwise implemented by, one or more server computing devices. In instances where server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0079] As described above, the server computing system 130 may store or otherwise include one or more text attribution models 140. For example, the model 140 may be or may otherwise include various machine learning models. Example machine learning models include neural networks or other multi-layer nonlinear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some example machine learning models may utilize attention mechanisms, such as self-attention. For example, some example machine learning models may include multi-head self-attention models (e.g., transformer models).
[0080] User computing device 102 and / or server computing system 130 may train models 120 and / or 140 via interaction with training computing system 150 communicatively coupled via network 180. Training computing system 150 may be separate from server computing system 130 or may be part of server computing system 130.
[0081] The training computing system 150 includes one or more processors 152 and memory 154. The one or more processors 152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or multiple processors operatively connected. The memory 154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 154 may store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes one or more server computing devices, or is otherwise implemented by one or more server computing devices.
[0082] The training computing system 150 may include a model trainer 160 that trains machine-learned models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques such as, for example, error back propagation. For example, a loss function may be back propagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques may be used to iteratively update parameters over multiple training iterations.
[0083] In some implementations, performing error backpropagation may include performing truncated backpropagation through time.The model trainer 160 may perform a variety of generalization techniques (eg, weight decay, dropout, etc.) to improve the generalization ability of the model being trained.
[0084] Specifically, model trainer 160 may train text attribution models 120 and / or 140 based on a set of training data 162 .
[0085] In some implementations, if the user has provided consent, the training examples may be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 may be trained by the training computing system 150 based on user-specific data received from the user computing device 102. In some instances, this process may be referred to as personalizing the model.
[0086] The model trainer 160 includes computer logic for providing the desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software that controls a general purpose processor. For example, in some implementations, the model trainer 160 includes a program file stored on a storage device, loaded into a memory, and executed by one or more processors. In other implementations, the model trainer 160 includes one or more computer executable instruction sets stored in a tangible computer readable storage medium such as RAM, a hard disk, or an optical or magnetic medium.
[0087] The network 180 may be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and may include any number of wired or wireless links. In general, communications over the network 180 may be conducted via any type of wired and / or wireless connection using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).
[0088] In some implementations, the input of the machine learning model of the present disclosure may be text or natural language data. The machine learning model may process text or natural language data to generate an output. As an example, the machine learning model may process natural language data to generate a language encoding output. As another example, the machine learning model may process text or natural language data to generate a potential text embedding output. As another example, the machine learning model may process text or natural language data to generate a translation output. As another example, the machine learning model may process text or natural language data to generate a classification output. As another example, the machine learning model may process text or natural language data to generate a text segmentation output. As another example, the machine learning model may process text or natural language data to generate a semantic intent output. As another example, the machine learning model may process text or natural language data to generate an upgraded (upscale) text or natural language output (e.g., text or natural language data with higher quality than the input text or natural language, etc.). As another example, the machine learning model may process text or natural language data to generate a prediction output.
[0089] In some implementations, the input of the machine learning model of the present disclosure may be speech data. The machine learning model may process speech data to generate an output. As an example, the machine learning model may process speech data to generate a speech recognition output. As another example, the machine learning model may process speech data to generate a speech conversion output. As another example, the machine learning model may process speech data to generate a potential embedding output. As another example, the machine learning model may process speech data to generate an encoded speech output (e.g., an encoded representation and / or a compressed representation of speech data, etc.). As another example, the machine learning model may process speech data to generate an upgraded speech output (e.g., speech data with a higher quality than the input speech data, etc.). As another example, the machine learning model may process speech data to generate a text representation output (e.g., a text representation of the input speech data, etc.). As another example, the machine learning model may process speech data to generate a predicted output.
[0090] In some implementations, the input of the machine-learned model of the present disclosure may be latent coded data (e.g., a latent space representation of an input, etc.). The machine-learned model may process the latent coded data to generate an output. As an example, the machine-learned model may process the latent coded data to generate a recognition output. As another example, the machine-learned model may process the latent coded data to generate a reconstruction output. As another example, the machine-learned model may process the latent coded data to generate a search output. As another example, the machine-learned model may process the latent coded data to generate a re-clustering output. As another example, the machine-learned model may process the latent coded data to generate a prediction output.
[0091] In some implementations, the input of the machine learning model of the present disclosure may be statistical data. Statistical data may be, represent, or otherwise include data calculated and / or computed from some other data source. The machine learning model may process statistical data to generate an output. As an example, the machine learning model may process statistical data to generate an identification output. As another example, the machine learning model may process statistical data to generate a prediction output. As another example, the machine learning model may process statistical data to generate a classification output. As another example, the machine learning model may process statistical data to generate a segmentation output. As another example, the machine learning model may process statistical data to generate a visualization output. As another example, the machine learning model may process statistical data to generate a diagnostic output.
[0092] In some implementations, the input of the machine learning model of the present disclosure may be sensor data. The machine learning model may process the sensor data to generate an output. As an example, the machine learning model may process the sensor data to generate a recognition output. As another example, the machine learning model may process the sensor data to generate a prediction output. As another example, the machine learning model may process the sensor data to generate a classification output. As another example, the machine learning model may process the sensor data to generate a segmentation output. As another example, the machine learning model may process the sensor data to generate a visualization output. As another example, the machine learning model may process the sensor data to generate a diagnostic output. As another example, the machine learning model may process the sensor data to generate a detection output.
[0093] In some cases, the machine learning model can be configured to perform a task that includes encoding input data for reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task can be an audio compression task. The input can include audio data, and the output can include compressed audio data. In another example, the input includes visual data (e.g., one or more images or videos), the output includes compressed visual data, and the task is a visual data compression task. In another example, the task can include generating an embedding for the input data (e.g., input audio or visual data).
[0094] In some cases, the input includes visual data, and the task is a computer vision task. In some cases, the input includes pixel data of one or more images, and the task is an image processing task. For example, the image processing task may be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the possibility that the one or more images depict an object belonging to the object class. The image processing task may be object detection, where the image processing output identifies one or more regions in the one or more images, and for each region, identifies the possibility that the region depicts an object of interest. As another example, the image processing task may be image segmentation, where the image processing output defines the corresponding possibility of each category in a set of predetermined categories for each pixel in the one or more images. For example, the group of categories may be foreground and background. As another example, the group of categories may be object classes. As another example, the image processing task may be depth estimation, where the image processing output defines a corresponding depth value for each pixel in the one or more images. As another example, the image processing task may be motion estimation, where the network input includes multiple images, and the image processing output defines the motion of the scene depicted at the pixel between the images in the network input for each pixel of one of the input images.
[0095] In some cases, the input includes audio data representing a spoken utterance, and the task is a speech recognition task. The output may include a text output mapped to the spoken utterance. In some cases, the task includes encrypting or decrypting the input data. In some cases, the task includes a microprocessor performance task, such as branch prediction or memory address translation.
[0096] Figure 1A An example computing system that can be used to implement the present disclosure is shown. Other computing systems may also be used. For example, in some implementations, the user computing device 102 may include a model trainer 160 and a training data set 162. In such implementations, the model 120 may be both trained and used locally at the user computing device 102. In some of such implementations, the user computing device 102 may implement the model trainer 160 to personalize the model 120 based on user-specific data.
[0097] Figure 1B A block diagram depicts an example computing device 1000 that performs in accordance with an example embodiment of the present disclosure. Computing device 1000 may be a user computing device or a server computing device.
[0098] Computing device 1000 includes multiple applications (e.g., applications 1 to N). Each application contains its own machine learning library and machine learning model. For example, each application can include a machine learning model. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc.
[0099] like Figure 1B As shown, each application can communicate with multiple other components of the computing device (such as, for example, one or more sensors, a context manager, a device state component, and / or additional components). In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to the application.
[0100] Figure 1C A block diagram depicts an example computing device 5000 that performs in accordance with an example embodiment of the present disclosure. Computing device 5000 may be a user computing device or a server computing device.
[0101] The computing device 5000 includes multiple applications (e.g., applications 1 to N). Each application communicates with a central intelligence layer. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc. In some implementations, each application can use an API (e.g., a public API across all applications) to communicate with the central intelligence layer (and the models stored therein).
[0102] The central intelligence layer includes multiple machine learning models. For example, Figure 1C As shown, a corresponding machine learning model can be provided for each application, and the corresponding machine learning model can be managed by the central intelligence layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model for all applications. In some implementations, the central intelligence layer is included in the operating system of the computing device 5000 or is otherwise implemented by the operating system.
[0103] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized data repository for the computing device 5000. Figure 1C As shown, the central device data layer can communicate with multiple other components of the computing device (such as, for example, one or more sensors, context managers, device state components, and / or additional components). In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0104] Figure 2 A block diagram depicting an example data workflow for editing and attributing machine-generated text according to an example embodiment of the present disclosure. Figure 2 In the example of the embodiment of the present invention, the first machine-learned language model 202 may output the first text content 204, and the query generation system 206 may generate one or more queries 208 related to aspects of the first text content 204. The evidence source retrieval system 209 may retrieve the evidence source 210 based on the query 208. Then, the evidence excerpt extraction system 212 may extract the evidence excerpt 214 from the evidence source 210. Then, the consistency / inconsistency detection system 216 may compare the extracted evidence excerpt 214 with the first text content 204 to determine one or more consistency indicators 218, which indicate consistency or inconsistency between the first text content 204 and one or more of the evidence excerpts 214. Then, the text editing system 220 may create an edited text 222 based on the first text content 204, the evidence excerpts 214, and the consistency indicators 218. The evidence attribution system 224 may attribute one or more aspects of the edited text 222 to the one or more evidence excerpts 214. Then, the evidence attribution system 224 may send the output 226 to the output system 228. Output 226 may include one or more texts (eg, first text content 204 , edited text 222 ) along with one or more attributions associating the evidence excerpt with the text.
[0105] Figure 2Depicts a first machine-learned language model 202 that outputs a first text content 204. The first machine-learned language model can be, for example, any machine-learned language model that is capable of generating text. In some instances, the first machine-learned language model 202 can be a multi-head self-attention model (e.g., a transformer). The first text content 204 can be, for example, natural language content. Although Figure 2 A first machine-learned model 202 is depicted outputting a first text content 204 , but those skilled in the art will appreciate that the text content may be generated or obtained in other ways, such as written by a human or retrieved from a database of text content.
[0106] Figure 2 A query generation system 206 is depicted that generates a query 208 based on a first text content 204. In some cases, the query generation system may include a machine learning model. In some instances, the query generation system may include a machine learning model using an attention mechanism, such as a multi-head self-attention model (e.g., a transformer). In some instances, a few-sample prompt may be used to enable a pre-trained general language model to perform query generation. In some instances, a pre-trained model may be prompted with a small number (e.g., six) of example input / output pairs, which have example text inputs and example associated query sets, wherein each example associated query set includes questions about each factual claim raised in each example text input. In some instances, the example input may also include an example context (e.g., "Context: Who is Abraham Lincoln? You said: He was the 16th president of the United States."). Once prompted in this manner, the pre-trained language model may be given a first text content 204 without a corresponding query set, and then the pre-trained model may generate a set of queries 208 as output. In some examples, the query generation system may be used multiple times (e.g., three times) to generate multiple query sets based on one first text content 204, and the query 208 may be a union of the multiple query sets. However, those skilled in the art will recognize that in some examples, a query may be generated without using a machine learning model (e.g., using the entire first text content 204 as a query, or parsing the first text content 204 into sentences and using each sentence as a query).
[0107] Figure 2Depicted is an evidence source retrieval system 209 that can retrieve evidence sources 210 based on the generated queries 208. In some instances, the evidence source retrieval system 209 can be a standard web search engine. For example, each of the generated queries 208 can be input into a standard web search engine, and a plurality (e.g., five) of the top-ranked results associated with each query can be added to the retrieved evidence sources 210. In other instances, the evidence source retrieval system 209 can retrieve evidence sources from a non-Internet-based corpus of evidence sources (e.g., a proprietary database containing proprietary data).
[0108] Figure 2 An evidence excerpt extraction system 212 is depicted that extracts evidence excerpts 214 from evidence sources 210. In some instances, the evidence excerpt extraction system 212 can extract candidate evidence excerpts from each web page by running a sliding window of four sentences (broken at the document title) across the evidence sources 210. In some instances, the candidate evidence excerpts can then be ranked based on their relevance to the corresponding query 208. In some instances, this ranking can be performed using a machine-learned model. In some instances, the ranking can be performed using a query-document relevance model that can calculate a relevance score between each query and each candidate evidence excerpt. In some cases, the query-document relevance model can include a text-to-text transformer suitable for transfer learning. In some instances, the query-document relevance model can include a pre-trained model that has been fine-tuned on a data set with relevance labels. Multiple (e.g., 5) top-ranked candidate evidence excerpts can then be selected for each query 208, and the evidence excerpt 214 can include the union of all top-ranked candidate evidence excerpts for all generated queries 208.
[0109] Figure 2Depicts a consistency / inconsistency detection system 216 that compares the extracted evidence excerpts 214 with the first text content 204 to determine one or more consistency indicators 218 that indicate consistency or inconsistency between the first text content 204 and one or more of the evidence excerpts 214. In some cases, the consistency / inconsistency detection system 216 may include a machine-learned model. In some instances, the consistency / inconsistency detection system 216 may include a machine-learned model that uses an attention mechanism, such as a multi-head self-attention model (e.g., a transformer). In some instances, a few-shot prompt may be used to enable a pre-trained general language model to perform consistency / inconsistency detection. In some instances, the few-shot prompt may also be a thought chain prompt, in which the pre-trained model is prompted to explicitly state an implicit answer to a query for each of the input text content and the input evidence excerpt. In some instances, a pre-trained model may be prompted with a small number (e.g., eight) of consistency check examples that include a text input, a query, an evidence excerpt, and an example output that explains whether the provided evidence excerpt and text input are consistent or inconsistent in answering the query. In some instances, the consistency check examples may also include example input context (e.g., "Context: Who is Abraham Lincoln?; You said: He was the 16th President of the United States.").
[0110] Figure 2Also depicted is a text editing system 220, which can then create an edited text 222 based on the first text content 204, the evidence excerpt 214, and the consistency indicator 218. In some instances, the text editing system 220 can be bypassed, and the first text content 204 can be output without editing, for example, if the consistency / inconsistency detection system 216 does not detect inconsistency. In some cases, the text editing system 220 may include a machine learning model. In some instances, the text editing system 220 may include a machine learning model using an attention mechanism, such as a multi-head self-attention model (e.g., a transformer). In some instances, a few sample prompts can be used to enable a pre-trained general language model to perform editing. In some instances, the few sample prompts can also be thought chain prompts, in which the pre-trained model is prompted to explicitly identify a section of text that needs to be edited in the first text content 204 before generating the edited text 222. In some instances, such thought chaining prompts can have the following benefits: enabling the text editing system 220 to retain a larger percentage of the first text content 204 while still correcting the errors identified in the first text content 204. In some instances, a pre-trained model can be prompted with a small number (e.g., seven) of examples having input text, a query, an evidence excerpt, a statement identifying a specific piece of text associated with an inconsistency between the evidence excerpt and the input text about the query, and an example edited input text. In some instances, the example input can also include an example context (e.g., "Context: Who is Abraham Lincoln?; You said: "He was the 16th President of the United States.").
[0111] Figure 2An evidence attribution system 224 that attributes one or more aspects of the edited text 222 to one or more evidence excerpts 214 is also depicted. In some instances, the evidence attribution system 224 may include a machine-learned model. In some instances, the machine-learned model used by the evidence attribution system 224 may also be used by the evidence excerpt extraction system 214. In some instances, the evidence attribution system 224 may include a query-document relevance model that can calculate a relevance score between each query and each evidence excerpt. In some cases, the query-document relevance model may include a text-to-text transformer suitable for transfer learning. In some instances, the query-document relevance model may include a pre-trained model that has been fine-tuned on a data set with relevance labels. Multiple (e.g., 5) evidence excerpts 214 may then be selected to create an attribution report that attributes the claims in the edited text 222 to one or more evidence excerpts. In some instances, evidence excerpts may be selected to maximize the overall coverage score of all generated queries 208. In some examples, the query-specific coverage score for a corresponding query 208 may be the maximum correlation score calculated between the query and all evidence excerpts included in the attribution report. In some examples, the overall coverage score may be the sum of the query-specific coverage scores for all generated queries 208. However, one skilled in the art will recognize that an attribution report may be compiled from evidence excerpts 214 in other ways. In some examples, the attribution report may be output separately from the edited text 222. In some examples, the attribution report may be integrated into the edited text 222 as an in-text reference.
[0112] Figure 2 Also depicted is an evidence attribution system 224 that sends an output 226 to an output system 228. The output 226 may include, for example, one or more texts (e.g., the first text content 204, the edited text 222) along with one or more attributions that associate an evidence excerpt with the text. The output system may be, for example, a display for display to a user, an output connection to another computing system, or an output connection to another computing program running on the same computing system as the evidence attribution system 224.
[0113] Example Method
[0114] Figure 3A block diagram depicting an example method for editing and / or attributing machine-generated text according to an example embodiment of the present disclosure. In step 302, a first text content 204 may be obtained. In step 304, one or more text documents containing one or more sources of evidence 210 may be retrieved based on the first text content 204. Optionally, in step 306, a second text content (e.g., edited text 222) may be generated based on the first text content 204 and the one or more text documents. Additionally or alternatively, in step 308, the first text content 204 or the second text content may be associated with one or more sources of evidence. In step 310, the first text content 204 or the second text content may be output, together with an attribution report associating one or more sources of evidence with one or more aspects of the first or second text content.
[0115] Step 302 describes obtaining the first text content 204. Those skilled in the art will recognize that the first text content 204 may be obtained in any suitable manner (e.g., using a machine learning model to generate the first text content; retrieving content from a database; receiving content as input from another computing system; etc.).
[0116] Step 304 depicts retrieving text documents containing one or more evidence sources 210 based on the first text content 204. In some examples, this step may include generating queries 208 using the query generation system 206, and then performing a standard web search using these queries. However, those skilled in the art will recognize that other methods of retrieving relevant text documents are also possible.
[0117] Step 306 depicts generating second text content based on the first text content 204 and the one or more text documents retrieved in step 304. In some instances, this step may include using the text editing system 220 to generate edited text 222. In some instances, this step may include using the consistency / inconsistency detection system 216 to determine one or more consistency indicators 218, and choosing not to edit if the consistency indicator 218 does not show an inconsistency. In other instances, the one or more consistency indicators 218 may be obtained in another manner, such as by receiving the consistency indicator 218 from another computing system or from a user.
[0118] Step 308 depicts associating the first text content 204 or the second text content with one or more evidence sources. In some cases, step 308 can include using the evidence attribution system 224 to create an attribution report associating the first text content 204 or the edited text 222 with the one or more evidence excerpts 214.
[0119] In step 310 , the first textual content 204 or the second textual content may be output, along with an attribution report associating one or more sources of evidence with one or more aspects of the first or second textual content. In some examples, step 310 may include outputting one or more output texts and attributions 226 using output system 228 .
[0120] Figure 4 A block diagram depicting an example method for retrieving evidence excerpts related to machine-generated text according to an example embodiment of the present disclosure. In step 402, a machine-learned model may identify one or more claims (e.g., factual claims, opinion claims) associated with a first text content 204. In some instances, step 402 may include using a query generation system 206. In some instances, step 402 may include a few-shot prompt or a chain of thought prompt of the query generation system 206. In step 404, a query 208 related to the identified claims may be generated using the machine-learned model. In some instances, step 404 may include using a query generation system 206. In some instances, step 404 may include a few-shot prompt or a chain of thought prompt of the query generation system 206. In some instances, steps 402 and 404 may be performed simultaneously (e.g., using a single chain of thought prompt of the query generation system 206). In other instances, step 402 may be bypassed or implicitly performed as part of step 404. In step 406, the generated query 208 may be used to retrieve the source of evidence 210. In some examples, step 408 can include using evidence source retrieval system 209. And in step 408, evidence excerpts 214 can be extracted from evidence sources 210 using a machine learning model. In some examples, step 408 can include using evidence excerpt extraction system 212.
[0121] Figure 5 A block diagram depicts an example method in which first text content 204 is compared to one or more retrieved evidence excerpts 214 according to an example embodiment of the present disclosure.
[0122] In step 502, the machine-learned language model may identify claims (e.g., factual claims, opinion claims) expressed or implied in the first text content 204. In some instances, step 502 may include using the consistency / inconsistency detection system 216. In some cases, step 502 may include a few-shot prompt or a thought chain prompt of the consistency / inconsistency detection system 216.
[0123] In step 504, the machine-learned model can identify claims expressed or implied in one or more evidence excerpts 214. In some instances, step 504 can include using consistency / inconsistency detection system 216. In some instances, step 504 can include a few-shot prompt or a chain of thought prompt from consistency / inconsistency detection system 216.
[0124] In step 506, the machine-learned model may identify a level of consistency or inconsistency between the claims associated with the first text content 204 and the claims associated with the evidence excerpt 214. In some instances, step 506 may include determining one or more consistency indicators 218 using the consistency / inconsistency detection system 216. In some instances, step 506 may include a few-shot prompt or a chain of thought prompt from the consistency / inconsistency detection system 216. In some instances, two or more of steps 502, 504, and 506 may be performed simultaneously (e.g., using a single chain of thought prompt from the consistency / inconsistency detection system 216). In other instances, one or more of steps 502 and 504 may be bypassed or performed implicitly as part of step 506.
[0125] In step 508, first text content 204 may optionally be edited to improve consistency between first text content 204 and evidence excerpt 214. In some instances, step 508 may include using text editing system 220 to generate edited text 222. In some instances, this step may include obtaining one or more consistency indicators 218, and choosing not to edit if consistency indicators 218 do not show inconsistencies. In other instances, one or more consistency indicators 218 may be obtained in another manner, such as by receiving consistency indicators 218 from another computing system or from a user.
[0126] In step 510, one or more evidence excerpts 214 may be associated with a claim expressed or implied in first textual content 204 or edited text 222 based on a level of consistency or inconsistency. In some cases, step 510 may include using evidence attribution system 224 to create an attribution report that associates first textual content 204 or edited text 222 with one or more evidence excerpts 214.
[0127] Figure 6A block diagram depicting an example method for editing machine-generated text according to an example embodiment of the present disclosure. In step 602, a computing system may identify a span of text to be edited. In some instances, step 602 may include using a text editing system 220. In some instances, step 602 may include a few sample prompts or a chain of thoughts prompt from the text editing system 220. In step 604, a computing system may generate one or more candidate edits. In some instances, step 604 may include using a text editing system 220. In some instances, step 604 may include a few sample prompts or a chain of thoughts prompt from the text editing system 220. In some instances, step 602 and step 604 may be performed simultaneously (e.g., using a single chain of thoughts prompt from the text editing system 220). In other instances, step 602 may be bypassed or implicitly performed as part of step 604.
[0128] In step 606, the computing system may determine whether the candidate edit improves the level of consistency between a series of evidence (e.g., one or more evidence excerpts 214) and the edited text content. In some instances, step 606 may include using the consistency / inconsistency detection system 216 on the edited text 222. In some instances, step 606 may include calculating an attribution metric that indicates the extent to which sentences in the edited text 222 are attributable to one or more evidence excerpts 214. In some instances, the attribution metric may include one or more attribution metrics disclosed in the example experiments section below.
[0129] In step 608, the computing system can determine whether the candidate edit preserves the content and structure of the first text content 204. In some examples, step 608 can include calculating a preservation score that indicates the degree to which relevant aspects of the first text content 204 are preserved by the edited text 222. In some examples, the preservation score can include one or more preservation metrics disclosed in the example experiments section below.
[0130] And in step 610, the computing system can select between the first text content 204 and the candidate edit. In some instances, step 610 can include selecting based on one or more of the retention score and the consistency score determined in steps 608 and 606, respectively. In some instances, step 610 can include selecting based on a combination of the retention score and the consistency score. In some instances, the combination can be based on a harmonic mean of the retention score and the consistency score.
[0131] Example Experiment
[0132] In this section, example metrics for measuring the degree to which a text can be attributed to one or more sources of evidence and for measuring the degree to which an edited text preserves the content of the original text are disclosed. Experiments according to the present disclosure are also disclosed. In the experiments, the systems and methods of the present disclosure improve the attribution of input text while better preserving the content and structure of the input text than previous work.
[0133] Regarding attribution, one available attribution measure is called Attributable to Identified Sources (AIS), which is a human evaluation framework that considers the binary concept of attribution. According to AIS, a text passage y is attributable to a set of evidence A if an average listener confirms the statement "According to A, y" based on the context of y. If everything in y can be attributed to A, the system gets full contribution (1.0), otherwise no contribution (0.0).
[0134] In some instances, a more fine-grained, sentence-level AIS extension can also be used. For example, annotators can be asked to give an AIS score for each sentence S of y, and then the average AIS score across all sentences is reported. Since the AIS score is binary, this effectively measures the percentage of sentences in y that are fully attributed to A.
[0135] When judging each sentence, annotators may also have access to one or more surrounding sentences and any other necessary context, such as the question that the text passage answers.
[0136] Automated metrics that approximate manual AIS judgments can also be used. For example, a machine-learned natural language inference (NLI) model can be used to estimate the model probability that the evidence excerpt 214 contains the corresponding sentence contained in the first text content 204 or the edited text 222. In some instances, the NLI model can be based on a text-to-text transformer model suitable for transfer learning. In some instances, the NLI model can be a pre-trained model that has been fine-tuned on the evidence-claim-implication triple, where the implication score can be "1" if the evidence implies the claim, and if the evidence does not imply the claim, the implication score can be "0". In some instances, the implication score associated with the evidence excerpt 214 and the sentence of the edited text 222 can be the model probability of outputting "1" given the evidence excerpt 214 and the sentence. For each sentence in the text input, multiple estimated implication probabilities can be calculated, for example, once for each of the multiple evidence excerpts. For each sentence, the maximum implication probability of all evidence excerpts can be the approximate AIS score of the sentence. Once a set of sentence-by-sentence AIS scores are approximated, an approximate AIS score for the entire text can be calculated by averaging the sentence-by-sentence scores. In some example experiments according to the present disclosure, the automatic attribution metric of the present disclosure has a strong Pearson correlation with a metric based on human evaluation.
[0137] In some instances, accuracy can be improved by decontextualizing each sentence based on the entire context of the text input.
[0138] With respect to preservation, an intent preservation metric can be generated by asking human annotators to decide whether a revision preserves the original intent of the text (e.g., completely, somewhat, or not at all). As with AIS evaluation, the annotators can be given any necessary surrounding context. A binary metric defining intent preservation can be 1.0 if the revision completely preserves the original intent and 0.0 otherwise.
[0139] In addition, a structure-preserving metric can be defined to measure whether the editing system has made unnecessary modifications, such as reordering words, changing the style of the text, or including unnecessary extra information. Different tasks may have different requirements for what should be preserved. In some instances, a metric that penalizes unnecessary changes can be easily calculated. For example, a structure-preserving metric can be based on the character-level Levenshtein edit distance between the input text and the output text (i.e., the minimum number of characters that must be changed or added in order to convert the input text to the output text). For example, the structure preservation metric can be equal to the greater of zero (0.0) and one minus the quotient calculated by dividing the Levenshtein edit distance by the input length characterized by the number of characters in the input text (1 - (lev(input, output) / length(input)). Such a metric can be 1.0 if the input and output are identical, and can be 0.0 if the output completely overwrites all of the input. The structure preservation metric can be calculated in this way without the use of any human annotators, which makes it useful for fully automated calculation of preservation metrics. Those skilled in the art will recognize that other metrics can be calculated and that specialized preservation metrics may be useful for specialized tasks (e.g., preserving rhyme schemes or puns). In some example experiments according to the present disclosure, the machine-calculated structure preservation metric of the present disclosure has a strong Pearson correlation with the preservation metric based on human evaluation.
[0140] A combined preservation metric may be calculated by combining the intent preservation metric and the structure preservation metric. For example, the combined preservation metric may be equal to the product calculated by multiplying the intent preservation metric by the structure preservation metric. Alternatively, the structure preservation metric or the intent preservation metric may be used alone or in combination with other metrics.
[0141] Additionally, the retention metric can be combined with the attribution metric to create a combined retention-attribution metric. For example, a harmonic mean of retention and attribution can be calculated. In some instances, this can be calculated in the same manner as the F1 score is calculated using the harmonic mean of precision and recall. The combined retention-attribution metric can be used to measure tasks that wish to maximize both attribution and retention while weighing any trade-offs between the two.
[0142] An advantage of the example retention and attribution metrics of the present disclosure is that no "golden" or "reference" edits are required (unlike many previous evaluations of text revision models), which are typically only available for specialized domains. This enables expansion of the scope to a wider range of generation tasks.
[0143] Figure 7 A block diagram depicting an example experiment and corresponding example output according to an example embodiment of the present disclosure, in which machine-generated text is edited and attributed to an evidence excerpt. Figure 7 , the text generation system 702 can generate the first text content 704. In some example experiments according to the present disclosure, the text generation system can be a 540 billion parameter decoder transformer-only based language model. In some instances, the first text content 704 can be generated by prompting the text generation system 702 with inputs designed to cause the language model to produce factual statements, reasoning chains, or knowledge-intensive dialogues. The depicted first text content 704 includes factual statements. Then, the research and revision system 706 can retrieve the evidence source 210 from the evidence source corpus 708, and can generate the second text content 710 and the attribution report 712. In some example experiments according to the present disclosure, the research and revision system 706 may include the query generation system 206, the evidence source retrieval system 209, the evidence excerpt extraction system 212, the consistency / inconsistency detection system 216, the text editing system 220, and the evidence attribution system 224. In some instances, the evidence source corpus 708 may include Internet-based evidence sources. The second textual content 710 and the attribution report 712 may then be provided to an output system 714. The depicted attribution report 712 includes two evidence excerpts 214.
[0144] Figure 8 A block diagram depicting an example experiment and corresponding example output according to an example embodiment of the present disclosure, in which machine-generated text is edited. Figure 8, the machine-generated text 802 can be input to a query generation system 804, which can then generate multiple queries 806 based on the machine-generated text 802. In some example experiments according to the present disclosure, the machine-generated text 802 can be generated by prompting a large (e.g., 540 billion parameters) decoder-only transformer-based language model with input designed to cause the language model to produce factual statements, reasoning chains, or knowledge-intensive dialogues. In some instances, the resulting paragraphs will be generally coherent but often contain factual errors. When prompted with input designed to produce knowledge-intensive dialogues, the machine-generated text 802 can be a context-dependent output characterized by pronouns and implicit references. In some instances, the query generation system 804 can include the query generation system 206.
[0145] The retrieval system 808 may use the query 806 to retrieve one or more evidence excerpts 810. The consistency system 812 may then compare the first evidence excerpt 810 to the machine-generated text 802. The consistency system 812 may output a consistency indicator 814 showing the inconsistency. When the inconsistency 814 is detected, the editing system 816 may edit the machine-generated text 802 to create an edited text 818. The consistency system 812 may then compare the second evidence excerpt 810 to the edited text 818, outputting a consistency indicator 820 showing the consistency. When a consistency 820 is detected, the edited text 818 may be sent to an output system 822 or consistency system without any additional editing to be compared with additional evidence excerpts 810. In some examples, the retrieval system 808, the consistency system 812, and the editing system 816 may include the evidence source retrieval system 209, the consistency / inconsistency detection system 216, and the text editing system 220, respectively.
[0146] In some example experiments according to the present disclosure, few-shot hints of a general large language model are used for the query generation step, the consistency / inconsistency detection step, and the editing step. Fig. 9Depicted are examples of few-shot prompts used to prompt an example query generation model 206, an example consistency / inconsistency detection system 216, and an example text editing system 220 in an example experiment according to the present disclosure. Block (a) depicts a query generation prompt 910, which includes an example input 902 and a corresponding example query set 904. Block (b) depicts an inconsistency detection prompt 912, which includes an example input 902, an example query 905, an example evidence excerpt 906, and an example inconsistency indicator 907. Block (c) depicts an example editing model prompt 914, which includes an example input 902, an example evidence query 905, an example evidence excerpt 906, an example inconsistency indicator 907, and an example edited text content 908.
[0147] In some example experiments according to the present disclosure, six query generation prompts 910 can be input into a large (e.g., 540 billion parameters) large language model based on a decoder-only transformer. In other example experiments, other model sizes (e.g., 62 billion parameters) can be used. This small amount of manually drafted examples can be sufficient for the pre-trained machine learning model to fully learn the task. In order to increase diversity and coverage, the machine learning model can be used three times to generate three sets of queries, and the generated queries 208 can include the union of the resulting queries. For each of the generated queries 208, a web search can be used to retrieve 5 web pages per query. In some example experiments, queries can be generated without using any machine learning query generation model (e.g., by using the entire first text content 204 as a query, or by using each sentence in the first text content 204 as a query). In some instances, experiments using machine learning query generation models will achieve a higher combined attribution-retention score than experiments using queries generated without machine learning models. Additionally, in some instances, the query generated by the machine-learned model may cause the evidence source retrieval system 209 to retrieve newer sources of evidence, or to retrieve relevant sources that are more likely to contradict an error in the first text input 204 .
[0148] In some example experiments according to the present disclosure, a query-document relevance model may be trained to compute relevance scores between queries 208 and evidence excerpts 214. Candidate evidence excerpts may then be extracted from each web page by running a sliding window of four sentences (broken at the document title) across the evidence source 210. One or more candidate evidence excerpts may then be ranked for each query based on one or more relevance scores between the query and each candidate evidence excerpt. In some example experiments, the highest scoring evidence excerpt for each query may be retained. In other instances, a different number of evidence excerpts may be retained.
[0149] In some example experiments according to the present disclosure, a small number of inconsistency detection prompts 912 can be input into a large (e.g., 540 billion parameters) decoder-only transformer-based large language model. In other example experiments, other model sizes (e.g., 62 billion parameters) can be used. In some instances, eight inconsistency detection prompts 912 can be used, which include example input 902, example query 905, example evidence excerpt 906, and example inconsistency indicators 907. This small number of manually drafted examples can be sufficient for a pre-trained machine learning model to fully learn the task. After the model has learned the task, the model can then be prompted with one or more input combinations including text input (e.g., first text content 204), generated query 208, and extracted evidence excerpt 214, so that the model outputs one or more consistency indicators 218.
[0150] In some example experiments according to the present disclosure, the text editing system 220 can process each first text content 204 without first detecting inconsistencies using any consistency / inconsistency detection system 216. In some instances, experiments using a machine-learned consistency detection model can achieve a higher combined attribution-retention score than experiments not using a consistency detection model.
[0151] In some example experiments according to the present disclosure, a small number of editing model prompts 914 can be input into a large language model based on a decoder-only transformer of 540 billion parameters. In other example experiments, other model sizes (e.g., 62 billion parameters) can be used. In some instances, the machine learning model can be prompted with seven editing model prompts 914, which include example input 902, example evidence query 905, example evidence excerpt 906, example inconsistency indicator 907, and example edited text content 908. This small number of manually drafted examples can be sufficient for the pre-trained machine learning model to fully learn the task. After the model has learned the task, the model can then be prompted with one or more input combinations including text input (e.g., first text content 204), generated query 208, extracted evidence excerpt 214, and consistency indicator 218, so that the model can output one or more edited texts 222.
[0152] In some example experiments according to the present disclosure, up to five evidence excerpts 214 may be selected to form an attribution report. In some instances, a query-document relevance model may be used to calculate a relevance score between each query 208 and each evidence excerpt 214. A set of evidence excerpts may then be selected to maximize coverage for all queries, where the coverage is calculated as the sum of query-specific coverage scores for each query 208 in all queries, where the query-specific coverage score is the maximum of a plurality of relevance scores calculated between the query 208 and each evidence excerpt 214 included in the attribution report.
[0153] In some example experiments according to the present disclosure, the outputs from the example systems were scored for retention and attribution, and these outputs were compared with the outputs from other systems. In some instances, the example embodiments of the present disclosure improve the attribution score of the edited text 222 relative to the first text content 204, while also maintaining the high retention score of the edited text 222 relative to the first text content 204. The example systems and methods of the present disclosure achieve attribution scores similar to other works, while achieving much higher retention scores, thereby achieving much higher combined attribution-retention scores. In some example environments, the systems and methods of the present disclosure retain the original intent of the input text 90% of the time, while other tested systems retain the original intent in 6% to 40% of the time. In some example environments, compared to unedited text content, the editing system according to the present disclosure increases attribution absolutely by up to 13%, while only changing 10% to 20% of the text.
[0154] Example Model Training
[0155] Fig.10 Flowchart depicting a method 1000 for training one or more machine-learned models according to aspects of the present disclosure. For example, an example machine-learned model may include a text generation model, such as a large language model.
[0156] One or more portions of the example method 1000 may be implemented by a computing system including one or more computing devices (such as, for example, the computing systems described with reference to other figures). Each respective portion of the example method 1000 may be performed by any one (or any combination) of the one or more computing devices. In addition, one or more portions of the example method 1000 may be implemented on hardware components of the apparatus described herein, for example, to train one or more systems or models. Fig.10 Elements performed in a specific order are depicted for purposes of illustration and discussion. One of ordinary skill in the art will appreciate, using the disclosure provided herein, that the elements of any method discussed herein may be adjusted, rearranged, expanded, omitted, combined or modified in various ways without departing from the scope of the present disclosure. Fig.10 Elements / terms described with reference to other systems and figures are described for illustrative purposes and are not intended to be limiting. One or more portions of the example method 1000 may additionally or alternatively be performed by other systems.
[0157] At 1002, example method 1000 may include obtaining training instances. A set of training data may include multiple training instances divided between multiple data sets (e.g., a training data set, a validation data set, or a test data set). The training instances may be labeled or unlabeled. Although referred to as "training" instances in example method 1000, it should be understood that when the model is trained (e.g., trained / learned online) using an evaluation of the performance of the model on the runtime instance, the runtime inference may form the training instance. Example data types for training instances and various tasks associated therewith are described throughout this disclosure.
[0158] At 1004, the example method 1000 may include processing the training instance using one or more machine-learned models to generate an output. The output may be obtained directly from the one or more machine-learned models, or may be a downstream result of a chain of processing operations that includes the output of the one or more machine-learned models.
[0159] At 1006, the example method 1000 may include receiving an evaluation signal associated with the output. The evaluation signal may be obtained using a loss function. Various losses may be determined, such as mean square error, likelihood loss, cross entropy loss, hinge loss, contrastive loss, or various other loss functions. The evaluation signal may be calculated using known baseline true labels (e.g., supervised learning), predicted or estimated labels (e.g., semi-supervised learning or self-supervised learning), or without labels (e.g., unsupervised learning). The evaluation signal may be a reward (e.g., for reinforcement learning). The reward may be calculated using a reward model of a machine learning configured to generate a reward based on the received output. The reward may be calculated using feedback data describing human feedback to the output.
[0160] At 1008, the example method 1000 may include updating the machine-learned model using the evaluation signal. For example, in some embodiments, various training or learning techniques (such as, for example, back propagation) may be used to learn the values of the parameters of the machine-learned model. For example, the evaluation signal may be back-propagated from the output (or another source of the evaluation signal) through the machine-learned model to update one or more parameters of the model (e.g., based on the gradient of the evaluation signal relative to the parameter value). For example, a system including one or more machine-learned models may be trained in an end-to-end manner. Gradient descent techniques may be used to iteratively update parameters in multiple training iterations. In some implementations, performing error back propagation may include performing truncated back propagation through time. The example method 1000 may include implementing a variety of generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.
[0161] In some implementations, example method 1000 can be implemented to train a machine learning model from an initialized state to a fully trained state (e.g., when the model exhibits a desired performance profile, such as based on accuracy, precision, recall, etc.).
[0162] In some implementations, the example method 1000 may be implemented for a specific stage of the training process. For example, in some implementations, the example method 1000 may be implemented for pre-training a machine learning model. Pre-training may include, for example, large-scale training on potentially noisy data to achieve a wide range of performance level bases across a variety of tasks / data types. In some implementations, the example method 1000 may be implemented for fine-tuning a machine learning model. Fine-tuning may include, for example, smaller-scale training on higher quality (e.g., labeled, selected, etc.) data. Fine-tuning may affect all or part of the parameters of the machine learning model. For example, various parts of the machine learning model may be "frozen" for certain training stages. For example, parameters associated with the embedding space may be "frozen" during fine-tuning (e.g., to retain information learned from a wider domain than that present in the fine-tuning data set). Example fine-tuning methods include reinforcement learning. Reinforcement learning may be based on user feedback on model performance during use.
[0163] Example machine learning model
[0164] Fig.11 is a block diagram of an example processing flow for using a machine learning model 1 to process input 2 to generate output 3.
[0165] The machine learning model 1 may be or include one or more machine learning models or model components. An example machine learning model may include a neural network (e.g., a deep neural network). An example machine learning model may include a nonlinear model or a linear model. An example machine learning model may use other architectures to replace or supplement a neural network. An example machine learning model may include a decision tree-based model, a support vector machine, a hidden Markov model, a Bayesian network, a linear regression model, a k-means clustering model, and the like.
[0166] Example neural networks may include feedforward neural networks, recurrent neural networks (RNNs) (including recurrent neural networks based on long short-term memory (LSTM)), convolutional neural networks (CNNs), diffusion models, generative adversarial networks, or other forms of neural networks. Example neural networks may be deep neural networks. Some example machine learning models may utilize attention mechanisms, such as self-attention. For example, some example machine learning models may include multi-head self-attention models.
[0167] The machine-learned model 1 may include a single or multiple instances of the same model configured to operate on data from the input 2. The machine-learned model 1 may include an ensemble of different models that can collaboratively interact to process data from the input 2. For example, the machine-learned model 1 may adopt a mixed expert structure. For example, see Zhou et al. Mixture-of-Experts with Expert Choice Routing, arXiv:2202.09368v2 (October 14, 2022).
[0168] Input 2 may generally include or otherwise represent various types of data. Input 2 may include one type of data or many different types of data. Output 3 may be the same type of data or a different type of data compared to input 2. Output 3 may include one type of data or many different types of data.
[0169] Example data types for input 2 or output 3 include natural language text data, software code data (e.g., source code, object code, machine code, or any other form of computer-readable instructions or programming language), machine code data (e.g., binary code, assembly code, or other form of machine-readable instructions that can be directly executed by a central processing unit of a computer), assembly code data (e.g., a low-level programming language that uses symbolic representations of machine code instructions to program a processing unit), genetic data or other chemical or biochemical data, image data, audio data, audio-visual data, tactile data, biometric data, medical data, financial data, statistical data, geographic data, astronomical data, historical data, sensor data in general (e.g., digital or analog values, such as voltage or other absolute or relative level measurements from real or artificial inputs (such as from audio sensors, light sensors, displacement sensors, etc.)), etc. The data can be raw or processed and can be in any format or mode.
[0170] Example machine learning sequence processing model
[0171] Fig.12 is a block diagram of an example implementation of an example machine learning model configured to process a sequence of information. For example, an example implementation of a machine learning model 1 may include a machine learning sequence processing model 4. An example system may pass an input 2 to the sequence processing model 4. The sequence processing model 4 may include one or more machine learning components. The sequence processing model 4 may process data from the input 2 to obtain an input sequence 5. The input sequence 5 may include one or more input elements 5-1, 5-2, ..., 5-M, etc. obtained from the input 2. The sequence processing model 4 may process the input sequence 5 using a prediction layer 6 to generate an output sequence 7. The output sequence 7 may include one or more output elements 7-1, 7-2, ..., 7-N, etc. generated based on the input sequence 5. The system may generate an output 3 based on the output sequence 7.
[0172] The sequence processing model 4 may include one or more machine learning model components configured to ingest, generate, or otherwise reason about sequences of information. For example, some example sequence processing models in the text domain are referred to as "large language models" or LLMs. See, for example, the PaLM 2 Technical Report, Google, https: / / ai.google / static / documents / palm2techreport.pdf (nd). Other example sequence processing models can operate in other domains, such as, for example, image domains, e.g., see Dosovitskiy et al., An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, arXiv:2010.11929v2 (June 3, 2021); audio domains, e.g., see Agostinelli et al., MusicLM: Generating Music From Text, arXiv:2301.11325v1 (January 26, 2023); biochemical domains, e.g., see Jumper et al., Highly accurate protein structure prediction with AlphaFold, 596 Nature 583 (August 26, 2021). Sequence processing models 4 can process one or more types of data simultaneously. The sequence processing model 4 may include a relatively larger model (eg, more parameters, computationally intensive, etc.), a relatively smaller model (eg, fewer parameters, computationally lightweight, etc.), or both.
[0173] In general, the sequence processing model 4 can use the data from the input 2 to obtain the input sequence 5. For example, the input sequence 5 can include a representation of the data from the input 2 in a format understood by the sequence processing model 4. One or more machine learning components of the sequence processing model 4 can ingest the data from the input 2, parse the data into segments compatible with the processing architecture of the sequence processing model 4 (e.g., via "lemmaization"), and project the segments into an input space associated with the prediction layer 6 (e.g., via "embedding").
[0174] The sequence processing model 4 may ingest data from the input 2 and parse the data into a sequence of elements to obtain an input sequence 5. For example, a portion of the input data from the input 2 may be decomposed into fragments that together represent the content of the portion of the input data. The fragments may provide elements of the sequence.
[0175] In some cases, elements 5-1, 5-2, ..., 5-M may represent building blocks for capturing or expressing meaningful information in a particular data domain. For example, an element may describe an "atomic unit" across one or more domains. For example, for a text input source, an element may correspond to a group of one or more word or sub-word components (such as a set of one or more characters).
[0176] For example, elements 5-1, 5-2, ..., 5-M may represent tokens obtained using a token analyzer. For example, a token analyzer may process a given portion of an input source and output a series of tokens representing that portion of the input source (e.g., corresponding to input elements 5-1, 5-2, ..., 5-M). Various methods may be used to tokenize. For example, a text input source may be tokenized using a byte pair encoding (BPE) technique. See, for example, Kudo et al., SentencePiece: A Simple and Language Independent Subword Tokenizer and Detokenizer for Neural Text Processing, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (System Demonstrations), pp. 66–71 (October 31–November 4, 2018), https: / / aclanthology.org / D18-2012.pdf. Image-based input sources can be tokenized by extracting and serializing patches from the image.
[0177] In general, any data type can be serialized and processed into an input sequence 5. It should be understood that Fig.12 The elements 5-1, 5-2, ..., 5-M depicted in FIG. 5 may be word-grams or may be embedded representations thereof.
[0178] The prediction layer 6 can predict one or more output elements 7-1, 7-2, ..., 7-N based on the input elements. The prediction layer 6 can include one or more machine learning model architectures, such as one or more learning parameter layers, which manipulate and transform the input to extract high-order meaning from the input elements 5-1, 5-2, ..., 5-M and the relationships between these input elements. In this way, for example, the example prediction layer 6 can predict new output elements based on the context provided by the input sequence 5.
[0179] The prediction layer 6 can evaluate associations between parts of the input sequence 5 and specific output elements. These associations can inform predictions about the likelihood that a specific output follows the input context. For example, consider the text snippet "The carpenter's toolbox is small and heavy. It is filled with ___." The example prediction layer 6 can recognize that "it" mentions "toolbox" again by determining the relationship between the corresponding embeddings. The example prediction layer 6 can also relate "it" to properties of the toolbox, such as "small" and "heavy." Based on these associations, for example, the prediction layer 6 can assign a higher probability to the word "nails" than to the word "sawdust."
[0180] A transformer is an example architecture that can be used in the prediction layer 4. For example, see Vaswani et al., AttentionIs All You Need, arXiv:1706.03762v7 (August 2, 2023). A transformer is an example of a model architecture for machine learning that uses an attention mechanism to compute associations between items within a context window. The context window may include a sequence that contains an input sequence 5 and possibly one or more output elements 7-1, 7-2, ..., 7-N. A transformer block may include one or more attention layers and one or more post-attention layers (e.g., feed-forward layers such as multi-layer perceptrons).
[0181] The prediction layer 6 may include other machine learning model architectures that complement or replace the transformer-based architecture. For example, recurrent neural networks (RNN) and long short-term memory (LSTM) models, as well as convolutional neural networks (CNNs), may also be used. In general, the prediction layer 6 may utilize various artificial neural networks that can understand or generate information sequences.
[0182] The output sequence 7 may include or otherwise represent the same or different data types as the input sequence 5. For example, the input sequence 5 may represent text data, and the output sequence 7 may represent text data. The input sequence 5 may represent image, audio, or audio-visual data, and the output sequence 7 may represent text data (e.g., describing the image, audio, or audio-visual data). It should be understood that the prediction layer 6 and any other gap model components of the sequence processing model 4 can be configured to receive multiple data types in the input sequence 5 and output multiple data types in the output sequence 7.
[0183] The output sequence 7 may have various relationships with the input sequence 5. The output sequence 7 may be a continuation of the input sequence 5. The output sequence 7 may be complementary to the input sequence 5. The output sequence 7 may translate, transform, enhance, or otherwise modify the input sequence 5. The output sequence 7 may answer, evaluate, confirm, or otherwise respond to the input sequence 5. The output sequence 7 may implement instructions provided via the input sequence 5 (or describe instructions for implementing the instructions).
[0184] The output sequence 7 can be generated autoregressively. For example, for some applications, the output of one or more prediction layers 6 can be passed through one or more output layers (e.g., a softmax layer) to obtain a probability distribution of an output vocabulary (e.g., a text or symbol vocabulary) conditioned on the set of input elements in the context window. In this way, for example, the output sequence 7 can be generated autoregressively by sampling a possible next output element, adding the element to the context window, and regenerating the probability distribution based on the updated context window, and sampling a possible next output element, etc.
[0185] The output sequence 7 can also be generated non-autoregressively. For example, multiple output elements of the output sequence 7 can be predicted together without explicit order conditions with respect to each other. See, for example, Saharia et al., Non-Autoregressive Machine Translation with Latent Alignments, arXiv:2004.07437v3 (16 November 2020).
[0186] The output sequence 7 may include one or more parts or elements. In an example content generation configuration, the output sequence 7 may include multiple elements corresponding to multiple parts of the generated output sequence (e.g., text sentences, values of discrete waveforms, computer code, etc.). In an example classification configuration, the output sequence 7 may include a single element associated with a classification output. For example, the output "vocabulary" may include a set of classes into which the input sequence will be classified. For example, the visual transformer block may pass the latent state information to a multilayer perceptron that outputs possible class values associated with the input image.
[0187] Fig.13 is a block diagram of an example technique for populating an example input sequence 8. The input sequence 8 may include various functional elements that form part of the model infrastructure, such as an element 8-0 obtained from a task indicator 9, which signals any model that processes the input sequence 8 that a specific task is being performed (e.g., to help adapt the performance of the model to the specific task). The input sequence 8 may include various data elements from different data modalities. For example, the input modality 10-1 may include a data modality. The data-to-sequence model 11-1 may process data from the input modality 10-1 to project the data into a format compatible with the input sequence 8 (e.g., one or more vectors of dimensions determined according to the dimensions of the input sequence 8) to obtain elements 8-1, 8-2, 8-3. Another input modality 10-2 may include different data modalities. The data-to-sequence model 11-2 may project the data from the input modality 10-2 into a format compatible with the input sequence 8 to obtain elements 8-4, 8-5, 8-6. Another input modality 10-3 may include yet another different data modality. The data-to-sequence model 11-3 may project the data from the input modality 10-3 into a format compatible with the input sequence 8 to obtain elements 8-7, 8-8, 8-9.
[0188] Input sequence 8 may be the same as or different from input sequence 5. Input sequence 8 may be a multimodal input sequence that includes elements that represent data from different modalities using a common dimensional representation. For example, the embedding space may have P dimensions. Input sequence 8 may be configured to include multiple elements having P dimensions. In this way, for example, example implementations may facilitate information extraction and reasoning across different data modalities by projecting data into elements in the same embedding space to perform comparisons, combinations, or other calculations between them.
[0189] For example, elements 8-0, ..., 8-9 may indicate specific locations within a multidimensional embedding space. Some elements may be mapped to a discrete set of locations in the embedding space. For example, elements corresponding to discrete members in a predetermined word-gram vocabulary may be mapped to discrete locations associated with those word-grams in the embedding space. Other elements may be distributed continuously across the embedding space. For example, some data types may be decomposed into continuously defined parts (e.g., image patches), which may be described using continuously distributed locations within the embedding space.
[0190] In some implementations, the expressive power of the embedding space may not be limited to the meaning associated with any particular set of word-grams or other building blocks. For example, a continuous embedding space can encode a series of high-order information. Individual fragments of information (e.g., word-grams) can be mapped to specific points in the space: for example, the word-gram of the word "dog" can be projected to an embedding value that points to a specific position in the embedding space associated with canine-related information. Similarly, an image patch of an image of a dog on grass can also be projected into the embedding space. In some implementations, the projection of the image of the dog can be similar to the projection of the word "dog", while also being similar to the projection of the word "grass", but potentially different from both. In some implementations, the projection of the image patch may not be completely aligned with any single projection of a single word. In some implementations, the projection of the image patch may be aligned with a combination of the projections of the words "dog" and "grass". In this way, for example, a high-order embedding space can encode information that may be independent of the data modality of the expressed information.
[0191] The task indicator 9 may include a model or model component configured to identify the task being performed and inject an input value represented by an element 8-0 into the input sequence 8, which element 8-0 signals which task is being performed. For example, the input value may be provided as a data type associated with the input modality and projected along with the input modality (e.g., the input value may be a text task label embedded along with other text data in the input; the input value may be a pixel-based representation of the task embedded along with other image data in the input; etc.). The input value may be provided as a data type that is different from or at least independent of other inputs. For example, the input value represented by the element 8-0 may be learned within a continuous embedding space.
[0192] The input modalities 10 - 1 , 10 - 2 , and 10 - 3 may be associated with a variety of different data types (eg, as described above with respect to input 2 and output 3 ).
[0193] The data-to-sequence models 11-1, 11-2, and 11-3 may be the same or different from each other. The data-to-sequence models 11-1, 11-2, and 11-3 may be adapted to each respective input modality 10-1, 10-2, and 10-3. For example, a text data-to-sequence model may subdivide a portion of the input text and project these subdivisions into elements in the input sequence 8 (e.g., elements 8-1, 8-2, 8-3, etc.). An image data-to-sequence model may subdivide an input image and project these subdivisions into elements in the input sequence 8 (e.g., elements 8-4, 8-5, 8-6, etc.). An arbitrary data type data-to-sequence model may subdivide the input of the arbitrary data type and project these subdivisions into elements in the input sequence 8 (e.g., elements 8-7, 8-8, 8-9, etc.).
[0194] The data-to-sequence models 11-1, 11-2, and 11-3 may form part of the machine-learned sequence processing model 4. The data-to-sequence models 11-1, 11-2, and 11-3 may be trained jointly with the machine-learned sequence processing model 4 or independently of the machine-learned sequence processing model 4. The data-to-sequence models 11-1, 11-2, and 11-3 may be trained end-to-end with the machine-learned sequence processing model 4.
[0195] Model development platform for sample machine learning
[0196] Fig.14 1 is a block diagram of an example model development platform 12 that can facilitate the creation, adaptation, and refinement of example machine learning models (e.g., machine learning model 1, sequence processing model 4, etc.). The model development platform 12 can provide a number of different tool kits that a developer system can use to develop new or adapted machine learning models.
[0197] The model development platform 12 may provide one or more model libraries 13 containing building blocks for new models. The model library 13 may include one or more pre-trained base models 13-1, which may provide a backbone of processing capabilities across various tasks. The model library 13 may include one or more pre-trained expert models 13-2, which may focus on performance in a specific professional field. The model library 13 may include various model primitives 13-3, which may provide low-level architectures or components (optionally pre-trained), which may be assembled in various arrangements as needed.
[0198] The model development platform 12 may receive a selection of various model components 14. The model development platform 12 may pass the selected model components 14 to a workbench 15, which combines the selected model components 14 into a development model 16.
[0199] The workbench 15 may facilitate further refinement and adaptation of the development model 16 by utilizing a number of different tool suites integrated with the model development platform 12. For example, the workbench 15 may facilitate alignment of the development model 16 with desired performance profiles on various tasks using a model alignment tool suite 17.
[0200] The model alignment toolkit 17 can provide a variety of tools for causing the development model 16 to generate outputs that are aligned with desired behavioral characteristics. Alignment can include increasing the accuracy, precision, recall, etc. of the model output. Alignment can include enforcing the use of output styles, patterns, or other preferred characteristics of the model output. Alignment can be general or domain-specific. For example, the pre-trained base model 13-1 can start from an initial performance level across multiple domains. Alignment of the pre-trained base model 13-1 can include improving performance in a specific information or task domain (e.g., even at the expense of performance in another information or task domain).
[0201] The model alignment toolkit 17 can integrate one or more datasets 17-1 for aligning the development model 16. The curated dataset 17-1 can include labeled or unlabeled training data. The dataset 17-1 can be obtained from a public domain dataset. The dataset 17-1 can be obtained from a private dataset associated with one or more developer systems for alignment of custom machine learning models customized for private use cases.
[0202] The pre-training pipeline 17-2 may include a model training workflow for machine learning configured to update the development model 16 on a large-scale, potentially noisy dataset. For example, pre-training may utilize unsupervised learning techniques (e.g., denoising, etc.) to process a large number of training instances to update model parameters from an initialized state and achieve a desired baseline performance. The pre-training pipeline 17-2 may utilize an unlabeled dataset in the dataset 17-1 to perform pre-training. The workbench 15 may implement the pre-training pipeline 17-2 to pre-train the development model 16.
[0203] The fine-tuning pipeline 17-3 may include a model training workflow of machine learning configured to use higher quality data to refine model parameters of the development model 16. The fine-tuning pipeline 17-3 may update the development model 16 by performing supervised training using a labeled dataset in the dataset 17-1. The fine-tuning pipeline 17-3 may update the development model 16 by performing reinforcement learning using a reward signal from a user feedback signal. The workbench 15 may implement the fine-tuning pipeline 17-3 to fine-tune the development model 16.
[0204] The prompt library 17-4 may include a set of inputs configured to induce behavior aligned with a desired performance criterion. The prompt library 17-4 may include few-shot prompts (e.g., inputs that provide examples of desired model outputs so that heads are appended to desired runtime queries), thought chain prompts (e.g., inputs that provide step-by-step reasoning within examples to encourage the model to perform comprehensive reasoning), and the like.
[0205] Example prompts may be retrieved from the available repository of prompt library 17-4. One or more developer systems may use workbench 15 to facilitate example prompts.
[0206] In some implementations, a pre-trained or fine-tuned model can achieve satisfactory performance without examples in the input. For example, a zero-trial prompt can include an input that lacks examples. The zero-trial prompt can be within the domain within the training dataset or outside the training domain.
[0207] The prompt library 17-4 may include one or more prompt engineering tools. The prompt engineering tool may provide a workflow for retrieving or learning optimized prompt values. The prompt engineering tool may facilitate directly learning prompt values (e.g., input element values) based on one or more training iterations. The workbench 15 may implement the prompt engineering tool in the development model 16.
[0208] The prompt library 17-4 may include a pipeline for prompt generation. For example, the development model 16 itself or another machine-learned model may be used to generate inputs. In this way, for example, a first model may process information about a task and output inputs for a second model to process in order to perform the steps of the task. The second model may be the same as or different from the first model. The workbench 15 may implement the prompt generation pipeline in the development model 16.
[0209] The prompt library 17-4 may include a pipeline for context injection. For example, the performance of the development model 16 on a particular task may be improved if provided with additional context for performing the task. The prompt library 17-4 may include software components configured to identify the desired context, retrieve the context from an external source (e.g., a database, a sensor, etc.), and add the context to the input prompt. The workbench 15 may implement the context injection pipeline in the development model 16.
[0210] Although various training examples described herein with respect to the model development platform 12 involve "pre-training" and "fine tuning", it should be understood that the model alignment toolkit 17 can generally support a variety of training techniques suitable for training a variety of machine learning models. The example training techniques can correspond to the example training method 1000 described above.
[0211] The model development platform 12 may include a model plug-in tool kit 18. The model plug-in tool kit 18 may include a variety of tools that are configured to enhance the functionality of the machine learning model by integrating the machine learning model with other systems, devices, and software components. For example, the machine learning model can use tools to improve the quality of performance where appropriate. For example, deterministic tasks can be offloaded to a dedicated tool instead of performing tasks in a probabilistic manner when the risk of error increases. For example, the machine learning model does not autoregressively predict the solution of the system of equations, but can identify the tools called to obtain the solution and pass the system of equations to the appropriate tool. The tool can be a traditional system of equations solver that can operate deterministically to solve the system of equations. The output of the tool can be returned in response to the original query. In this way, the use of tools can allow some example models to focus on the advantages of machine learning models-for example, understanding the intent in unstructured requests for tasks-while enhancing the performance of the model by offloading certain tasks to more focused tools to mechanically apply deterministic algorithms to well-defined problems.
[0212] The model plug-in tool suite 18 may include a validation tool 18-1. The validation tool 18-1 may include a tool that can parse and validate the output of a machine-learned model. The validation tool 18-1 may include an engineered heuristic that establishes certain thresholds to be applied to the model output. For example, the validation tool 18-1 may ground the output of the machine-learned model with a structured data source (e.g., to mitigate "hallucinations").
[0213] The model plug-in toolkit 18 may include a toolkit 18-2 for implementing one or more tools, which may include scripts or other executable code that can be executed with the development model 16. The toolkit 18-2 may include one or more inputs configured to enable the machine learning model to implement the tool (e.g., a few sample prompts that induce the model to output tool calls with the correct syntax, etc.). For example, the toolkit 18-2 may include fine-tuning training data for training the model to use the tool.
[0214] The model plug-in tool suite 18 may include an interface for calling an external application programming interface (API) 18-3. For example, in addition to or instead of directly implementing tool calls or tool code using the development model 16, the development model 16 may also be aligned with output instructions that initiate API calls to send or obtain data via an external system.
[0215] The model plug-in toolkit 18 may be integrated with the hint library 17-4 to build a catalog of available tools for use with the development model 16. For example, the model may receive in input a catalog of available tools, and the model may generate an output that selects a tool from the available tools and initiates a tool call for using the tool.
[0216] The model development platform 12 may include a computing optimization tool suite 19 for optimizing the computing performance of the development model 16. For example, a model compression 19-1 tool may allow the development model 16 to reduce size while maintaining a desired performance level. For example, the model compression 19-1 may include quantization workflows, weight pruning, and sparsification techniques, etc. The hardware acceleration 19-2 tool may facilitate the configuration of model storage and execution formats to operate optimally on different hardware resources. For example, the hardware acceleration 19-2 may include tools for optimally slicing the model for distributed processing on multiple processing units to increase bandwidth, reduce unified memory requirements, etc. The distillation 19-3 tool may be used to train a more lightweight model based on the knowledge encoded in the development model 16. For example, the development model 16 may be a high-performance large-scale machine learning model optimized using the model development platform 12. In order to obtain a lightweight model for running in a resource-constrained environment, the smaller model may be a "student model" that learns to imitate the development model 16 as a "teacher model". In this way, for example, the investment in learning the parameters and configuration of the development model 16 can be efficiently transferred to a smaller model for more efficient inference.
[0217] The workbench 15 may implement one, more than one, or no tool suite implemented in the model development platform 12. The workbench 15 may output an output model 20 based on the development model 16. The output model 20 may be a deployed version of the development model 16. The output model 20 may be a development or training checkpoint of the development model 16. The output model 20 may be a distilled, compressed, or otherwise optimized version of the development model 16.
[0218] Fig.15 1 is a block diagram of an example training process for training a development model 16 for machine learning. One or more portions of the example training process may be implemented by a computing system including one or more computing devices (such as, for example, the computing systems described with reference to other figures). Each respective portion of the example training process may be performed by any one (or any combination) of the one or more computing devices. In addition, one or more portions of the example training process may be implemented on hardware components of the devices described herein, for example, to train one or more systems or models. Fig.15Elements performed in a specific order are depicted for purposes of illustration and discussion. One of ordinary skill in the art will appreciate, using the disclosure provided herein, that the elements of any method discussed herein may be adjusted, rearranged, expanded, omitted, combined or modified in various ways without departing from the scope of the present disclosure. Fig.15 Elements / terms described with reference to other systems and figures are described for illustrative purposes and are not intended to be limiting. Additionally or alternatively, one or more portions of the example training process may be performed by other systems.
[0219] Initially, the development model 16 may be maintained in an initial state as an initialized model 21. The development model 16 may be initialized using weight values. The initial weight values may be random or based on an initialization pattern. The initial weight values may be based on previous pre-training for the same or a different model.
[0220] The initialized model 21 may be pretrained in a pretraining stage 22. The pretraining stage 22 may be implemented using one or more pretraining pipelines 17-2 on data from the dataset 17-1. For example, if the initialized model 21 has been pretrained (e.g., the development model 16 includes, is, or is based on a pretrained base model or expert model), pretraining may be omitted.
[0221] The pre-trained model 23 may then be a new version of the development model 16, which may be maintained as the development model 16 or a new development model. If the development model 16 has been pre-trained, the pre-trained model 23 may be the initial state. The pre-trained model 23 may be fine-tuned in a fine-tuning phase 24. The fine-tuning phase 24 may be implemented using one or more fine-tuning pipelines 17-3 on data from the dataset 17-1. For example, if the pre-trained model has satisfactory performance, if the model has already been fine-tuned, or if other methods of adjustment are preferred, fine-tuning may be omitted.
[0222] The fine-tuned model 29 may then be a new version of the development model 16, which may be maintained as the development model 16 or a new development model. If the development model 16 has already been fine-tuned, the fine-tuned model 29 may be the initial state. The fine-tuned model 29 may undergo refinement 26 with user feedback. For example, refinement 26 with user feedback may include performing reinforcement learning, optionally based on human feedback from human users of the fine-tuned model 25. Since reinforcement learning may be a form of fine-tuning, it should be understood that the fine-tuning stage 24 may include a stage for performing refinement 26 with user feedback. Refinement 26 with user feedback may produce a refined model 27. The refined model 27 may be output to a downstream system 28 for deployment or further development.
[0223] In some implementations, computational optimization operations may be applied before, during, or after each stage. For example, the initialized model 21 may be computationally optimized 29-1 (e.g., using the computational optimization toolkit 19) before the pre-training stage 22. The pre-trained model 23 may be computationally optimized 29-2 (e.g., using the computational optimization toolkit 19) before the fine-tuning stage 24. The fine-tuned model 25 may be computationally optimized 29-3 (e.g., using the computational optimization toolkit 19) before refinement 26 using user feedback. The refined model 27 may be computationally optimized 29-4 (e.g., using the computational optimization toolkit 19) before being output to the downstream system 28. The computational optimizations 29-1, ..., 29-4 may all be the same, all different, or include at least some different optimization techniques.
[0224] Example machine learning model inference system
[0225] Fig.16 1 is a block diagram of an inference system for operating one or more machine learning models 1 to perform inference (e.g., for training, for deployment, etc.). A model host 31 may receive a machine learning model 1. The model host 31 may host one or more model instances 31-1, which may be one or more instances of one or more models. The model host 31 may host the model instance 31-1 using available computing resources 31-2 associated with the model host 31.
[0226] Model master 31 may perform inference on behalf of one or more clients 32. Client 32 may transmit input request 33 to model master 31. Using input request 33, model master 31 may obtain input 2 to input into machine learning model 1. Machine learning model 1 may process input 2 to generate output 3. Using output 3, model master 31 may return output payload 34 in response to input request 33 from client 32. Output payload 34 may include or be based on output 3.
[0227] The model host 31 can utilize various other resources and tools to enhance the inference task. For example, the model host 31 can communicate with the tool interface 35 to facilitate the model instance 31-1 to use the tool. The tool interface 35 may include a local or remote API. The tool interface 35 may include an integrated script or other software function. The model host 31 may use an online learning interface 36 to promote the continuous improvement of the machine learning model 1. For example, the online learning interface 36 can be used in a reinforcement learning loop to retrieve user feedback on the inference served by the model host 31. The model host 31 can access a runtime data source 37 for enhancing the input 2 with additional context information. For example, the runtime data source 37 may include a knowledge graph 37-1 that facilitates structured information retrieval for information associated with the input request 33 (e.g., a search engine service). The runtime data source 37 may include a public or private, external or local database 37-2 that can store information associated with the input request 33 for enhancing the input 2. The runtime data source 37 may include account data 37 - 3 , which may be retrieved in association with a user account corresponding to the client 32 to customize the behavior of the model host 31 accordingly.
[0228] The model host 31 may be implemented by one or more computing devices or systems. The client 2 may be implemented by one or more computing devices or systems, which may include a computing device or system shared with the model host 31 .
[0229] For example, the model host 31 may be operated on a server system that provides machine learning services (e.g., over a local area network or wide area network) to client devices operating the client 32. The client device may be an end-user device used by an individual. The client device may be a server system that operates the client 32 to provide various functions as services to downstream end-user devices.
[0230] In some implementations, the model host 31 may operate on the same device or system as the client 32. The model host 31 may be a machine learning service that runs on a device to provide machine learning functionality to one or more applications operating on a client device, which may include an application that implements the client 32. The model host 31 and the client 32 may be part of the same application. For example, the model host 31 may be a subroutine or method implemented by a portion of the application, and the client 32 may be another subroutine or method that uses the model host 31 to perform inference functionality within the application. It should be understood that the model host 31 and the client 32 may have a variety of different configurations.
[0231] Model instance 31-1 may include one or more machine learning models that can be used to perform inference. Model instance 31-1 may include weights or other model components stored on / in a persistent storage device, temporarily cached, or loaded into a high-speed memory. Model instance 31-1 may include multiple instances of the same model (e.g., for executing more requests in parallel on the same model). Model instance 31-1 may include instances of different models. Model instance 31-1 may include cached intermediate states of active or inactive models, and cached intermediate states are used to accelerate the inference of those models. For example, an inference session with a particular model may generate a significant amount of computational results that can be reused for future inference runs (e.g., using a KV cache for a transformer-based model). These computational results may be stored in association with the inference session so that the session can be executed more efficiently when restored.
[0232] Computing resources 31-2 may include one or more processors (central processing units, graphics processing units, tensor processing units, machine learning accelerators, etc.) connected to one or more memory devices. Computing resources 31-2 may include a dynamic pool of available resources shared with other processes. Computing resources 31-2 may include a memory device large enough to fit the entire model instance in a single memory instance. Computing resources 31-2 may also share model instances across multiple memory devices (e.g., using data parallelism or tensor parallelism, etc.). Doing so can increase parallelism or execute large models using multiple memory devices that individually may not be able to fit the entire model in memory.
[0233] The input request 33 may include data for the input 2. The model host 31 may process the input request 33 to obtain the input 2. The input 2 may be obtained directly from the input request 33 or may be retrieved using the input request 33. The input request 33 may be submitted to the model host 31 via an API.
[0234] The model host 31 can perform inference on multiple batches of input requests 33 in parallel. For example, the model instance 31-1 can be configured with an input structure having a batch dimension. Individual inputs 2 can be distributed across the batch dimension (e.g., rows of an array). Individual inputs 2 can include completely different contexts. Individual inputs 2 can be multiple inference steps of the same task. Individual inputs 2 can be interleaved in the input structure so that any given inference cycle can operate on different parts of the corresponding input 2. In this way, for example, the model host 31 can perform inference on batches in parallel, so that the output 3 can also contain a batch dimension and return the inference results of the batched inputs 2 in parallel. In this way, for example, multiple batches of input requests 33 can be processed in parallel to achieve a higher throughput of the output payload 34.
[0235] The output payload 34 may include or be based on the output 3 from the machine-learned model 1. The model host 31 may process the output 3 to obtain the output payload 34. This may include chaining multiple rounds of inference (e.g., iteratively, recursively, across the same model or different models) to obtain the final output of the task to be returned in the output payload 34. The output payload 34 may be transmitted to the client 32 via an API.
[0236] The online learning interface 36 can facilitate reinforcement learning of the machine-learned model 1. The online learning interface 36 can facilitate reinforcement learning with human feedback (RLHF). The online learning interface 36 can facilitate federated learning of the machine-learned model 1.
[0237] The model host 31 can execute the machine learning model 1 to perform inference for various tasks using various types of data. For example, various different inputs 2 and outputs 3 can be used for various different tasks. In some implementations, the input 2 can be or otherwise represent image data. The machine learning model 1 can process the image data to generate an output. As an example, the machine learning model 1 can process the image data to generate an image recognition output (e.g., recognition of image data, potential embedding of image data, encoded representation of image data, hash of image data, etc.). As another example, the machine learning model 1 can process the image data to generate an image segmentation output. As another example, the machine learning model 1 can process the image data to generate an image classification output. As another example, the machine learning model 1 can process the image data to generate an image data modification output (e.g., a change in image data, etc.). As another example, the machine learning model 1 can process the image data to generate an encoded image data output (e.g., an encoded and / or compressed representation of image data, etc.). As another example, the machine learning model 1 can process the image data to generate an upgraded image data output. As another example, the machine learning model 1 can process the image data to generate a prediction output.
[0238] In some implementations, the task is a computer vision task. In some cases, the input 2 includes pixel data of one or more images, and the task is an image processing task. For example, the image processing task may be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the possibility that the one or more images depict an object belonging to the object class. The image processing task may be object detection, where the image processing output identifies one or more regions in the one or more images, and for each region, identifies the possibility that the region depicts an object of interest. As another example, the image processing task may be image segmentation, where the image processing output defines the corresponding possibility of each category in a set of predetermined categories for each pixel in the one or more images. For example, the group of categories may be foreground and background. As another example, the group of categories may be object classes. As another example, the image processing task may be depth estimation, where the image processing output defines a corresponding depth value for each pixel in the one or more images. As another example, the image processing task may be motion estimation, where the network input includes multiple images, and the image processing output defines the motion of the scene depicted at the pixel between images in the network input for each pixel of one of the input images.
[0239] In some implementations, input 2 may be or otherwise represent natural language data. The machine learning model 1 may process natural language data to generate an output. As an example, the machine learning model 1 may process natural language data to generate a language encoding output. As another example, the machine learning model 1 may process natural language data to generate a potential text embedding output. As another example, the machine learning model 1 may process natural language data to generate a translation output. As another example, the machine learning model 1 may process natural language data to generate a classification output. As another example, the machine learning model 1 may process natural language data to generate a text segmentation output. As another example, the machine learning model 1 may process natural language data to generate a semantic intent output. As another example, the machine learning model 1 may process natural language data to generate an upgraded text or natural language output (e.g., a text or natural language data with a higher quality than the input text or natural language, etc.). As another example, the machine learning model 1 may process natural language data to generate a prediction output (e.g., one or more predicted next parts of the natural language content).
[0240] In some implementations, input 2 may be or otherwise represent speech data (e.g., data describing spoken natural language, such as audio data, text data, etc.). The machine-learned model 1 may process speech data to generate an output. As an example, the machine-learned model 1 may process speech data to generate a speech recognition output. As another example, the machine-learned model 1 may process speech data to generate a speech translation output. As another example, the machine-learned model 1 may process speech data to generate a potential embedding output. As another example, the machine-learned model 1 may process speech data to generate an encoded speech output (e.g., an encoded and / or compressed representation of speech data, etc.). As another example, the machine-learned model 1 may process speech data to generate an upgraded speech output (e.g., speech data of higher quality than the input speech data, etc.). As another example, the machine-learned model 1 may process speech data to generate a text representation output (e.g., a text representation of the input speech data, etc.). As another example, the machine-learned model 1 may process speech data to generate a predicted output.
[0241] In some implementations, input 2 may be or otherwise represent potentially coded data (e.g., a latent space representation of an input, etc.). The machine-learned model 1 may process the potentially coded data to generate an output. As an example, the machine-learned model 1 may process the potentially coded data to generate a recognition output. As another example, the machine-learned model 1 may process the potentially coded data to generate a reconstruction output. As another example, the machine-learned model 1 may process the potentially coded data to generate a search output. As another example, the machine-learned model 1 may process the potentially coded data to generate a re-clustering output. As another example, the machine-learned model 1 may process the potentially coded data to generate a prediction output.
[0242] In some implementations, input 2 may be or otherwise represent statistical data. Statistical data may be, represent, or otherwise include data calculated and / or computed from some other data source. Machine-learned model 1 may process statistical data to generate an output. As an example, machine-learned model 1 may process statistical data to generate an identification output. As another example, machine-learned model 1 may process statistical data to generate a prediction output. As another example, machine-learned model 1 may process statistical data to generate a classification output. As another example, machine-learned model 1 may process statistical data to generate a segmentation output. As another example, machine-learned model 1 may process statistical data to generate a visualization output. As another example, machine-learned model 1 may process statistical data to generate a diagnostic output.
[0243] In some implementations, input 2 may be or otherwise represent sensor data. Machine-learned model 1 may process sensor data to generate an output. As an example, machine-learned model 1 may process sensor data to generate a recognition output. As another example, machine-learned model 1 may process sensor data to generate a prediction output. As another example, machine-learned model 1 may process sensor data to generate a classification output. As another example, machine-learned model 1 may process sensor data to generate a segmentation output. As another example, machine-learned model 1 may process sensor data to generate a visualization output. As another example, machine-learned model 1 may process sensor data to generate a diagnostic output. As another example, machine-learned model 1 may process sensor data to generate a detection output.
[0244] In some implementations, the machine learning model 1 can be configured to perform tasks including encoding input data to achieve reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task can be an audio compression task. The input can include audio data, and the output can include compressed audio data. In another example, the input includes visual data (e.g., one or more images or videos), the output includes compressed visual data, and the task is a visual data compression task. In another example, the task can include generating embeddings for input data (e.g., input audio or visual data). In some cases, the input includes audio data representing spoken utterances, and the task is a speech recognition task. The output can include text output mapped to spoken utterances. In some cases, the task includes encrypting or decrypting the input data. In some cases, the task includes microprocessor performance tasks, such as branch prediction or memory address conversion.
[0245] In some implementations, the task is a generative task, and the machine-learned model 1 can be configured to output content generated from input 2. For example, input 2 can be or otherwise represent data of one or more modalities that encodes context for generating additional content.
[0246] In some implementations, the task may be a text completion task. The machine-learned model 1 may be configured to process an input 2 representing text data and generate an output 3 representing additional text data that completes a text sequence including the input 2. For example, the machine-learned model 1 may be configured to generate an output 3 to complete a sentence, paragraph, or portion of text following a portion of text represented by the input 2.
[0247] In some implementations, the task may be an instruction-following task. The machine learning model 1 may be configured to process an input 2 representing an instruction for performing a function and generate an output 3 that advances the goal of satisfying the instruction function (e.g., at least one step of a multi-step process for performing the function). The output 3 may represent data of the same or different modality as the input 2. For example, the input 2 may represent text data (e.g., a natural language instruction for a task to be performed), and the machine learning model 1 may process the input 2 to generate an output 3 that represents text data in response to the instruction (e.g., a natural language response, a programming language response, a machine language response, etc.). The input 2 may represent image data (e.g., an image-based instruction for a task to be performed, optionally with a text instruction), and the machine learning model 1 may process the input 2 to generate an output 3 that represents text data in response to the instruction (e.g., a natural language response, a programming language response, a machine language response, etc.). One or more outputs 3 may be generated iteratively or recursively to sequentially process and complete steps toward completing the requested function. For example, the initial output can be executed by an external system or processed by a machine learning model 1 to complete the initial steps of executing the function. Multiple steps can be performed and a final output is obtained in response to the initial instruction.
[0248] In some implementations, the task may be a question-answering task. The machine learning model 1 may be configured to process an input 2 representing a question to be answered and generate an output 3 that advances the goal of returning an answer to the question (e.g., at least one step of a multi-step process for performing the function). The output 3 may represent data of the same or different modality as the input 2. For example, the input 2 may represent text data (e.g., natural language instructions for a task to be performed), and the machine learning model 1 may process the input 2 to generate an output 3 that represents text data in response to the question (e.g., a natural language response, a programming language response, a machine language response, etc.). The input 2 may represent image data (e.g., image-based instructions for a task to be performed, optionally with text instructions), and the machine learning model 1 may process the input 2 to generate an output 3 that represents text data in response to the question (e.g., a natural language response, a programming language response, a machine language response, etc.). One or more outputs 3 may be generated iteratively or recursively to sequentially process and complete steps toward answering the question. For example, the initial output can be executed by an external system or processed by the machine learning model 1 to complete the initial steps of obtaining the answer to the question (e.g., querying a database, performing calculations, executing a script, etc.). Multiple steps can be performed and a final output responsive to the question is obtained.
[0249] In some implementations, the task may be an image generation task. The machine learning model 1 may be configured to process an input 2 representing a context about a desired portion of the image content. The context may include text data, image data, audio data, etc. The machine learning model 1 may be configured to generate an output 3 representing image data depicting an image associated with the context. For example, the machine learning model 1 may be configured to generate pixel data for an image. The value of a channel associated with a pixel in the pixel data may be selected based on the context (e.g., based on a probability determined based on the context).
[0250] In some implementations, the task may be an audio generation task. The machine learning model 1 may be configured to process an input 2 representing a context about a desired portion of the audio content. The context may include text data, image data, audio data, etc. The machine learning model 1 may be configured to generate an output 3 representing audio data associated with the context. For example, the machine learning model 1 may be configured to generate waveform data in the form of an image (e.g., a spectrogram). The value of a channel associated with a pixel of the image may be selected based on the context. The machine learning model 1 may be configured to generate waveform data in the form of a sequence of discrete samples of a continuous waveform. The value of the sequence may be selected based on the context (e.g., based on a probability determined based on the context).
[0251] In some implementations, the task may be a data generation task. The machine learning model 1 may be configured to process an input 2 representing a context about an expected portion of data (e.g., data from various data domains, such as sensor data, image data, multimodal data, statistics, etc.). For example, the expected data may be synthetic data used to train other machine learning models. The context may include any data type. The machine learning model 1 may be configured to generate an output 3 representing data aligned with the expected data. For example, the machine learning model 1 may be configured to generate data values for populating a data set. The value of a data object may be selected based on the context (e.g., based on a probability determined based on the context).
[0252] Additional public content
[0253] The technology discussed herein relates to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for multiple possible configurations, combinations, and partitions of tasks and functions between and among components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system, or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0254] Although the present subject matter has been described in detail with respect to various specific example embodiments of the present subject matter, each example is provided by way of explanation rather than limitation of the present disclosure. Those skilled in the art may easily produce changes, variations, and equivalents to such embodiments after understanding the foregoing. Therefore, the present disclosure does not exclude such modifications, variations, and / or additions to the present subject matter that would be readily apparent to those of ordinary skill in the art. For example, a feature shown or described as part of one embodiment may be used together with another embodiment to produce yet another embodiment. Therefore, the present disclosure is intended to encompass such changes, variations, and equivalents.
Claims
1. A computer-implemented method for improving the attribution of content generated by a model, the method comprising: obtaining, by a computing system including one or more computing devices, first text content generated by a first machine-learned language model; retrieving, by the computing system, one or more sources of evidence related to the first text content generated by the first machine-learned language model; generating, by the computing system, second text content as an output of a second machine-learned language model based on a comparison between the first text content and the one or more sources of evidence; and associating, by the computing system, one or more portions of the one or more sources of evidence as an attribution of the first or second text content.
2. The method according to claim 1, wherein retrieving, by the computing system, the one or more sources of evidence comprises: generating, by the computing system, one or more queries based on the first text content; and retrieving, by the computing system, the one or more sources of evidence based on the one or more queries.
3. The method according to claim 2, wherein retrieving, by the computing system, the one or more sources of evidence based on the one or more queries includes a web search.
4. The method according to any one of the preceding claims, wherein generating, by the computing system, one or more queries based on the first text content includes prompting a third machine-learned language model.
5. The method according to any one of the preceding claims, wherein generating, by the computing system, one or more queries based on the first text content comprises: identifying one or more claims associated with the first text content; and generating a query associated with at least one of the one or more claims.
6. The method according to any one of the preceding claims, comprising: performing at least one of the following: prompting a third machine-learned language model to identify one or more claims associated with the first text content; or prompting a third machine-learned language model to generate a query associated with at least one of the one or more claims.
7. The method according to claim 6, wherein prompting the third machine-learned language model includes few-shot prompting.
8. The method according to claim 6 or 7, wherein prompting the third machine-learned language model includes chain-of-thought prompting.
9. The method according to any one of the preceding claims, further comprising: extracting, by the computing system, one or more evidence excerpts from the one or more sources of evidence retrieved by the computing system; and comparing, by the computing system, the one or more evidence excerpts extracted by the computing system with the first text content.
10. The method according to claim 9, wherein extracting, by the computing system, one or more evidence excerpts from the one or more sources of evidence retrieved by the computing system comprises: determining one or more candidate evidence excerpts; for each candidate evidence excerpt, generating a corresponding score based on the query used to retrieve the one or more sources of evidence; and selecting a subset of the candidate evidence excerpts based on the corresponding scores.
11. The method according to claim 10, wherein the respective scores are relevance scores generated by a query-document relevance model of machine learning.
12. The method according to claim 10 or 11, wherein: the one or more evidence sources are text documents; and determining one or more candidate evidence excerpts includes running a sliding window across the respective text documents.
13. The method according to claim 10, 11 or 12, wherein selecting a subset of the candidate evidence excerpts based on the respective scores includes: maximizing the coverage of the one or more claims.
14. The method according to any one of claims 9 to 12, wherein comparing the one or more evidence excerpts extracted by the computing system with the first text content by the computing system includes: determining one or more first consistency levels between the one or more evidence excerpts and the first text content with respect to one or more claims associated with the first text content.
15. The method according to claim 14, wherein determining one or more first consistency levels between the one or more evidence excerpts and the first text content with respect to one or more claims associated with the first text content includes: utilizing a consistency model of machine learning to evaluate the first text content and the evidence excerpts.
16. The method according to claim 15, wherein utilizing a consistency model of machine learning to evaluate the first text content and the evidence excerpts includes prompting the consistency model of machine learning with chain-of-thought prompting.
17. The method according to claim 15, wherein utilizing a consistency model of machine learning to evaluate the first text content and the evidence excerpts includes prompting the consistency model of machine learning with few-shot prompting.
18. The method according to any one of claims 14, 15, 16 and 17, wherein generating second text content by the computing system as an output of a second language model of machine learning includes: identifying a respective claim among the one or more claims associated with the first text content, wherein the first consistency level with respect to the respective claim indicates a lack of perfect consistency; and generating second text content such that a second consistency level between the second text content and the one or more evidence excerpts with respect to the respective claim indicates a higher degree of consistency than the first consistency level with respect to the respective claim.
19. The method according to any one of the preceding claims, wherein generating second text content by the computing system as an output of a second language model of machine learning includes using chain-of-thought prompting to prompt the second language model of machine learning.
20. The method according to any one of the preceding claims, wherein generating second text content by the computing system as an output of a second language model of machine learning includes using few-shot prompting to prompt the second language model of machine learning.
21. The method according to any one of the preceding claims, further includes: determining an edit distance based on one or more differences between the first text content and the second text content; and Determine whether to output the second text content based on the edit distance.
22. The method according to any one of the preceding claims, further comprising: Determine a consistency score between the second text content and one or more evidence excerpts extracted from the one or more evidence sources; Determine a retention score indicating the similarity between the first text content and the second text content; and Based on a combination of the consistency score and the retention score, determine whether to output the second text content.
23. The method according to any one of the preceding claims, wherein: The one or more evidence sources are text documents.
24. A computing system, comprising: One or more processors; and One or more non-transitory computer-readable media storing instructions that can be executed by the one or more processors to cause the computing system to perform one or more operations, the operations including: Obtain first text content generated by a first machine-learned language model; Retrieve one or more evidence sources related to the first text content generated by the first machine-learned language model; Generate second text content as an output of a second machine-learned language model based on a comparison between the first text content and the one or more evidence sources; and Associate one or more parts of the one or more evidence sources as an attribution of the first or second text content.
25. One or more computer-readable media storing instructions that can be executed by a computing system to cause the computing system to perform one or more operations, the operations comprising: Obtain first text content generated by a first machine-learned language model; Retrieve one or more evidence sources related to the first text content generated by the first machine-learned language model; Generate second text content as an output of a second machine-learned language model based on a comparison between the first text content and the one or more evidence sources; and Associate one or more parts of the one or more evidence sources as an attribution of the first or second text content.
Citation Information
Cited By
Verifiable text generation method and device
CN122019736A