Generating an output document via a machine-learned model based on a document template

CA3320508A1Pending Publication Date: 2025-08-14GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CA3320508
Authority / Receiving Office
CA · CA
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-09
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Current computing systems require significant effort and computational resources to create specific prompts for large language models (LLMs) to process textual content, leading to wasted time and resources due to the need for manual experimentation or retraining.

Method used

A computing system is configured to automatically extract style and structure from source documents and apply it to generate a document template, allowing users to create new documents efficiently by leveraging machine-learned models without manual prompting.

Benefits of technology

This approach reduces computational waste by automating the generation of document templates and summaries, saving time and resources by eliminating the need for manual prompting and retraining, while improving the accuracy and reliability of generated content.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A computing device for generating an output document includes one or more memories to store instructions and one or more processors to execute the instructions to perform operations, the operations including: receiving a plurality of source documents, receiving an input associated with a request to generate the output document based on the plurality of source documents and a particular document template, and generating, via one or more machine-learned models, the output document having the particular document template.
Need to check novelty before this filing date? Find Prior Art

Description

GENERATING AN OUTPUT DOCUMENT VIA A MACHINE-EEARNED MODELBASED ON A DOCUMENT TEMPLATEFIELD[1] The disclosure relates generally to generating content via one or more machine-learned models based on source content that is selected (identified) by a user. For example, the disclosure relates to methods and computing devices for generating content (an output document) by implementing a document template generated by one or more machine-learned models with respect to the source content, thereby assisting the user in managing content, organizing content, creating content, etc.BACKGROUND[2] According to current computing systems, large language models (LLMs) are capable of interacting with textual content. For example, a user may copy and paste content from one document into a chat box to query the LLM about the content. The LLM may provide an output (e.g., a summary) regarding the content.SUMMARY[3] Aspects and advantages of embodiments of the disclosure will be set forth in part in the following description, or can be learned from the description, or can be learned through practice of the example embodiments.[4] In one or more example embodiments, a computing device for generating, organizing, managing, and creating content is provided. For example, the computing device includes: one or more memories configured to store instructions; and one or more processors configured to execute the instructions to perform operations, the operations comprising: receiving a plurality of source documents, receiving an input associated with a request to generate the output document based on the plurality of source documents and a particular document template, and generating, via one or more machine-learned models, the output document having the particular document template.[5] In some implementations, the operations further comprise: providing, for presentation on a display device, a graphical user interface configured to receive a selection of the plurality of source documents.[6] In some implementations, the operations further comprise: determining, by the one or more machine-learned models, an intent associated with the output document, based on a content associated with each of the plurality of source documents.[7] In some implementations, the operations further comprise: determining, by the one or more machine-learned models, a style associated with the output document, based on a content of each of the plurality of source documents.[8] In some implementations, the operations further comprise: determining, by the one or more machine-learned models, a format associated with the output document, based on a content of each of the plurality of source documents.[9] In some implementations, determining, by the one or more machine-learned models, the format associated with the output document, based on the content of each of the plurality of source documents, includes identifying a plurality of headings for respective sections of the output document.

[0010] In some implementations, the operations further comprise: providing, for presentation on a display device, a graphical user interface comprising a plurality of user interface elements, wherein each of the plurality of user interface elements corresponds to a different style which can be applied by the one or more machine-learned models for generating the output document having the particular document template; and receiving a selection of one or more of the plurality of user interface elements, wherein generating, via the one or more machine-learned models, the output document having the particular document template comprises applying one or more styles corresponding to the selection of the one or more of the plurality of user interface elements.

[0011] In some implementations, the operations further comprise: providing, for presentation on a display device, a graphical user interface comprising a plurality of user interface elements, wherein each of the plurality of user interface elements corresponds to a different heading which can be applied by the one or more machine-learned models for generating sections of the output document having the particular document template; and receiving a selection of one or more of the plurality of user interface elements, wherein generating, via the one or more machine-learned models, the output document having the particular document template comprises applying one or more headings corresponding to the selection of the one or more of the plurality of user interface elements to sections of the output document.

[0012] In some implementations, the computing device further comprises one or more databases configured to store a plurality of generative machine-learned models respectively associated with a plurality of different document types, and the operations further comprise retrieving, from among the plurality of generative machine-learned models, at least onegenerative machine-learned model associated with a document type indicated by common content of the plurality of source documents.

[0013] In some implementations, the operations further comprise: receiving a plurality of training documents associated with a document type, wherein each of the plurality of training documents include one or more common features associated with the particular document template; and training the one or more machine-learned models to leam at least one of an intent, a sty le, or a format of the particular document template for the document type.

[0014] In one or more example embodiments, a computer-implemented method for organizing, managing, and creating content is provided. The computer-implemented method comprises receiving, by a computing system comprising one or more processors and one or more machine-learned models, a plurality of source documents; receiving, by the computing system, an input associated with a request to generate an output document based on the plurality of source documents and a particular document template; and generating, via the one or more machine-learned models, the output document having the particular document template.

[0015] In some implementations, the method further comprises determining, by the one or more machine-learned models, at least one of a style, an intent, and a format associated with the output document, based on a content of each of the plurality of source documents.

[0016] In some implementations, the method further comprises providing, for presentation on a display device of the computing system, a graphical user interface comprising a plurality of user interface elements, wherein each of the plurality of user interface elements corresponds to a sty le, an intent, or a format which can be applied by the one or more machine-learned models for generating the output document having the particular document template; and receiving a selection of one or more of the plurality of user interface elements, wherein generating, via the one or more machine-learned models, the output document having the particular document template comprises applying the style, the intent, or the format corresponding to the selection of the one or more of the plurality7of user interface elements.

[0017] In some implementations, the method further comprises storing, in one or more databases, a plurality of generative machine-learned models respectively associated with a plurality of different document types; and retrieving, from among the plurality of generative machine-learned models, at least one generative machine-learned model associated with a document type indicated by common content of the plurality of source documents.

[0018] In some implementations, the method further comprises receiving a plurality of training documents associated with a document type, wherein each of the plurality of trainingdocuments include one or more common features associated with the particular document template; and training the one or more machine-learned models to learn at least one of an intent, a style, or a format of the particular document template for the document type.

[0019] In one or more example embodiments, a computing device for generating, organizing, managing, and creating content is provided. For example, the computing device includes: one or more memories configured to store instructions; and one or more processors configured to execute the instructions to perform operations, the operations comprising: receiving a selection of a plurality of sample documents; receiving an input to generate the document template based on the selection of the plurality of sample documents; and in response to receiving the input, implementing one or more machine-learned models to generate the document template based on at least one of a style, a format, or an intent associated with the plurality of sample documents.

[0020] In some implementations, the operations further comprise: receiving a plurality of source documents, receiving a further input associated with a request to generate an output document based on the plurality of source documents and the document template, and generating, via the one or more machine-learned models, the output document according to the document template.

[0021] In some implementations, the operations further comprise: determining, by the one or more machine-learned models, the at least one of the style, the format, and the intent associated with the plurality of sample documents, based on a content and structure of each of the plurality of sample documents.

[0022] In some implementations, the operations further comprise: providing, for presentation on a display device, a graphical user interface comprising a plurality of user interface elements, wherein at least one of the plurality of user interface elements is selectable to change the at least one of the style, the format, and the intent associated with the plurality of sample documents determined by the one or more machine-learned models; and receiving a selection of one or more of the plurality of user interface elements to change the at least one of the style, the format, and the intent associated with the plurality of sample documents determined by the one or more machine-learned models, wherein generating, via the one or more machine-learned models, the output document according to the document template comprises applying the at least one of the style, the format, and the intent changed according to the selection of the one or more of the plurality of user interface elements.

[0023] In some implementations, the computing device further comprises one or more databases configured to store a plurality of generative machine-learned models respectivelyassociated with a plurality of different document types, and the operations further comprise retrieving, from among the plurality of generative machine-learned models, at least one generative machine-learned model associated with a document type indicated by common content of the plurality of sample documents.

[0024] In one or more example embodiments, a computer-implemented method for organizing, managing, and creating content is provided. The computer-implemented method comprises receiving, by a computing system comprising one or more processors, a selection of a plurality of sample documents; receiving, by the computing system, an input to generate the document template based on the selection of the plurality of sample documents; and in response to receiving the input, implementing one or more machine-learned models of the computing system, to generate the document template based on at least one of a style, a format, or an intent associated with the plurality of sample documents.

[0025] In one or more example embodiments, a computer-readable medium (e.g., a non- transitory computer-readable medium) which stores instructions that are executable by one or more processors of a computing system is provided. In some implementations the computer- readable medium stores instructions which may include instructions to cause the one or more processors to perform one or more operations which are associated with any of the methods described herein (e.g., operations of the server computing system and / or operations of the computing device). The computer-readable medium may store additional instructions to execute other aspects of the server computing system and computing device and corresponding methods of operation, as described herein.

[0026] These and other features, aspects, and advantages of various embodiments of the disclosure will become better understood with reference to the following description, drawings, and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the disclosure and, together with the description, serve to explain the related principles.BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Detailed discussion of example embodiments directed to one of ordinary skill in the art is set forth in the specification, which makes reference to the appended drawings, in which:

[0028] FIGS. 1 A-1B depict example systems according to according to one or more example embodiments of the disclosure;

[0029] FIG. 2 illustrates a flow diagram of an example, non-limiting computer-implemented method, according to one or more example embodiments of the disclosure;

[0030] FIG. 3 depicts an example block diagram of a notebook application, according to one or more example embodiments of the disclosure;

[0031] FIGS. 4A-4H illustrate example user interface screens of a notebook application, according to one or more example embodiments of the disclosure;

[0032] FIGS. 5A-5B illustrate example user interface screens of a notebook application, according to one or more example embodiments of the disclosure;

[0033] FIGS. 6A-6B illustrate further example user interface screens of a notebook application, according to one or more example embodiments of the disclosure;

[0034] FIG. 7 illustrates example notebooks or projects which can be represented in a particular manner, according to one or more example embodiments of the disclosure;

[0035] FIGS. 8A-8B illustrate flow diagrams of example, non-limiting computer-implemented methods, according to one or more example embodiments of the disclosure;

[0036] FIG. 9 depicts an example block diagram of a document extractor application, according to one or more example embodiments of the disclosure;

[0037] FIGS. 10A-10E illustrate example actions associated with a document extractor application, according to one or more example embodiments of the disclosure;

[0038] FIG. 11 A depicts a block diagram of an example computing system for organizing, managing, and creating content by implementing one or more machine-learned models with respect to training content and / or source content selected by a user, according to one or more example embodiments of the disclosure;

[0039] FIG. 1 IB depicts a block diagram of an example computing device for organizing, managing, and creating content by implementing one or more machine-learned models with respect to training content and / or source content selected by a user, according to one or more example embodiments of the disclosure;

[0040] FIG. 11C depicts a block diagram of an example computing device for organizing, managing, and creating content by implementing one or more machine-learned models with respect to training content and / or source content selected by a user, according to one or more example embodiments of the disclosure.DETAILED DESCRIPTION

[0041] Reference now will be made to embodiments of the disclosure, one or more examples of which are illustrated in the drawings, wherein like reference characters denote like elements. Each example is provided by way of explanation of the disclosure and is not intended to limit the disclosure. In fact, it will be apparent to those skilled in the art thatvarious modifications and variations can be made to disclosure without departing from the scope or spirit of the disclosure. For instance, features illustrated or described as part of one embodiment can be used with another embodiment to yield a still further embodiment. Thus, it is intended that the disclosure covers such modifications and variations as come within the scope of the appended claims and their equivalents.

[0042] Terms used herein are used to describe the example embodiments and are not intended to limit and / or restrict the disclosure. The singular forms ‘"a,” "‘an7’ and "the ' are intended to include the plural forms as well, unless the context clearly indicates otherwise. In this disclosure, terms such as "including", "having", “comprising”, and the like are used to specify features, numbers, steps, operations, elements, components, or combinations thereof, but do not preclude the presence or addition of one or more of the features, elements, steps, operations, elements, components, or combinations thereof.

[0043] It will be understood that, although the terms first, second, third, etc., may be used herein to describe various elements, the elements are not limited by these terms. Instead, these terms are used to distinguish one element from another element. For example, without departing from the scope of the disclosure, a first element may be termed as a second element, and a second element may be termed as a first element.

[0044] The term "and / or" includes a combination of a plurality of related listed items or any item of the plurality of related listed items. For example, the scope of the expression or phrase "A and / or B" includes the item "A", the item "B", and the combination of items "A and B”.

[0045] In addition, the scope of the expression or phrase "at least one of A or B" is intended to include all of the following: (1) at least one of A, (2) at least one of B, and (3) at least one of A and at least one of B. Likewise, the scope of the expression or phrase "at least one of A. B, or C" is intended to include all of the following: (1) at least one of A, (2) at least one of B, (3) at least one of C, (4) at least one of A and at least one of B, (5) at least one of A and at least one of C, (6) at least one of B and at least one of C, and (7) at least one of A, at least one of B, and at least one of C.

[0046] Large Language Models (LLMs) are capable of generating generic textual content. However, if a user wants to leverage an LLM to generate a specific type or style of content, the user either needs to experiment w ith a number of different prompting strategies, or retrain (finetune) the LLM on a number of training samples that demonstrate the specific type or style of content. Experimenting with different prompting strategies can result in unnecessary and redundant processing / content generation as the user will try a number of times to createthe desired content before the appropriate sty le is achieved. Retraining the model is highly expensive in terms of computational usage. Therefore, both of these approaches result in wasted computational resources.

[0047] According to examples of the disclosure, a computing system is configured to automatically extract a ty pe or style from one or more source documents (e.g., notes) and then apply the extracted type or style to generate a template for creating a new (output) document. For example, a computing system may be trained to learn a particular template based on a plurality of documents of a particular type. As an example, a user can upload a plurality7of product requirements documents (PRDs) and the computing system may be configured to leam the document structure and style. Subsequently, the user (or another user) can upload a plurality of source documents (e.g., a plurality of user experience research (UXR) documents) and provide a request for a document to be generated having a particular format (e.g., a PRD format). The computing system may be configured to generate the PRD based on the plurality' of source documents and based on the learned document structure and style.

[0048] In some implementations, the user can be provided with a graphical user interface to change or modify a style, format, and / or intent of the document template. For example, if the document template has a document format including a title section, description section, background section, and conclusion section which are provided in a particular order, the user may be provided with a graphical user interface which includes a plurality of user interface elements that can be selected to modify an order or arrangement of the sections. For example, if the document template has a particular sty le (e.g., opinionated), the user may be provided with a graphical user interface which includes a plurality' of user interface elements that can be selected to change the sty le from the document template to another style (e.g., casual language).

[0049] In some implementations, the computing system can be configured to train one or more machine-learned models to leam attributes of a plurality' of training documents to leam a document type and / or to leam an intent, a style, and a format of the plurality of training documents. The plurality of training documents may share one or more common features (e.g., a same or similar document type, a same or similar document structure, a same or similar style, etc.).

[0050] In some implementations, the computing system includes one or more databases configured to store a plurality of generative machine-learned models respectively associated with a plurality7of different document types. The computing system may be configured toretrieve, from among the plurality of generative machine-learned models, at least one generative machine-learned model associated with a document type indicated by common content of the plurality of source documents. For example, the plurality of source documents may each contain content which is associated with a particular type of document (e.g., a resume, a PRD, a research paper, etc.).

[0051] According to current computing systems, large language models (LLMs) are capable of interacting with textual content. However, current computing systems require significant effort to create a specific prompt for a LLM to process. For example, a user may be required to copy and paste content from one document into a chat box to query the LLM about the content. This switching between multiple windows or applications results in significant amounts of wasted computational time and resources (e.g., processor cycles).

[0052] According to examples of the disclosure, a computing system (computing platform, computing device) is configured to create a new type of output (e.g., an outline, a report, a summary, etc.) via one or more machine-learned models, based on source content provided to the computing system (e g., by the user). For example, the computing system may be configured to receive source content selected by a user and generate, via one or more machine-learned models, a summary of the source content including an identification of one or more topics related to the source content.

[0053] As an example, a user may identify and select a subset of documents (e.g., four documents) from a plurality of documents (a large corpus of documents) relating to a topic (e.g., modem American history in the 1990s) which are provided to the computing system. The computing system may include one or more machine-learned models configured to receive as an input the selected documents and to provide as an output a summary' (or a report, a paper, an outline, etc.) relating to the selected documents and an identification of key topics (e.g., via a document guide).

[0054] In some implementations, the computing system is configured to implement a semantic retrieval method (e.g., clustering) and one or more machine-learned models (e.g., one or more LLMs) to generate a summary, key topics, and suggested queries (e.g., questions) to produce a document guide for content identified or indicated by the user (e.g.. based on a body of text found in the content).

[0055] For example, in some implementations the computing system may be configured to receive source content from the user. For example, the user may upload source content (e.g., documents, imagery, sound files, websites, videos, presentations, PDFs, etc.). In some implementations, the computing system may be configured to, in response to the useruploading source content, automatically generate information including a summary of the source content, generate top themes found in the source content, generate suggested topics and questions to help the user explore the source content further, etc. The information may be presented via a user interface. The user interface may be configured to receive an input from the user (e.g., via a touch-input, mouse-click, etc.) on a user interface element corresponding to a theme, question, etc.. In response to receiving the input from the user, the computing system may be configured to respond to the input, for example, by providing an answer via one or more machine-learned models to the question or theme query, based on the source content.

[0056] In some implementations, the computing system may be configured to, in response to the user uploading source content, automatically generate information including a report, an outline, or a rewrite of the original content, so as to generate new content based on the source content identified (selected) by the user. For example, the user may request that the computing system identify a specified number of themes from one or more documents, to summarize client interactions occurring over a specified duration of time (e.g., the least two weeks), to generate a specified number of ideas based on a source document, etc.

[0057] In some implementations, the source content that is relied upon or referenced by the LLMs may be selected (e.g., curated) by the user. For example, the user may consider or indicate that the selected source content is trustworthy (e.g., trusted source content, authoritative source content, etc.) or has a higher priority compared to other content which does not have such a designation. Therefore, the one or more machine-learned models are configured to generate summaries of content, or generate new content, based on trusted source content, improving the accuracy and reliability of information and data provided to the user. Further, the one or more machine-learned models are configured to answer questions about the source content based on the trusted source content, improving the accuracy and reliability' of information and data provided as answers to questions posed by the user.

[0058] In some implementations, the computing system can be configured to discover, add, or remove source content. For example, the user may add or remove source content. For example, the user may provide an input requesting the computing system to discover source content (e g., by conducting a search for scholarly articles regarding a certain topic) and the user may add the discovered source content as part of the selected source content which is deemed trustworthy by the user (and / or the computing system).

[0059] In some implementations, the computing system may be configured to receive an additional source content by the user creating a new note, by the user uploading the sourcecontent to the computing system, by adding the source content via a website, etc. The computing system may be configured to generate or receive metadata concerning the added source content. For example, the metadata may include one or more of a title, an author, a date of upload, a date associated with the creation of the source content, a uniform resource location (URL) associated with the source content, etc.

[0060] In some implementations, the computing system may be configured to delete or remove a source content by the user selecting the source content and providing an input requesting that the source content be deleted (e.g., from the notepad application). In some implementations, the source content may be deleted as a source relied upon for generating summaries, key topics, etc., in the notepad application, but an original copy of the source content may be maintained elsewhere.

[0061] In some implementations, the computing system can be configured to receive a user input via a text entry box (e.g., an open-ended text entry' box). For example, the user input may be in the form of a question (e.g., “What did Nixon say in his speech about automobile use”). For example, the user input may be in the form of a theme or idea (e.g., “Nixon automobile crisis” or “What is this document about?”).

[0062] The computing system may be configured to, via one or more machine-learned models, provide a response to the user input based on the selected source content. In some implementations, the computing system is configured to indicate the number of sources (citations) that were relied upon for providing the response. In some implementations, the computing system is configured to provide for presentation a source (citation) which was relied upon for a particular passage in the response. In some implementations, the computing system is configured to provide additional context regarding the source (citation) which was relied upon for the particular passage in the response. For example, the computing system may indicate the passage (e.g., a sentence or paragraph) from the source for which a portion of the response was based on and may further indicate a preceding and / or subsequent passage from the source to provide further context concerning the particular passage.

[0063] In some implementations, the computing system may be configured to store one or more passages (e.g., snippets) from a generated response (answer) to a query (question) input by the user. For example, the one or more passages may be stored in a specified area of a notepad application. The specified area may be referred to as a scratchpad and each item of information stored in the scratchpad may be referred to as a note. The one or more passages may be selected by the user for storing as a first note in the scratchpad. In some implementations, citations can be stored in the scratchpad as a second note. In someimplementations, the user can select (e.g., highlight) a particular passage from a citation (source content) for storing in the scratchpad as a third note. In some implementations, the user can store their own passage or comments as a written note (fourth note).

[0064] According to examples of the disclosure, the computing system may be configured to provide a notepad application which is configured to generate an output (e.g., an outline, a report, a summary, etc.) via one or more machine-learned models, based on source content provided to the notebook application (e.g., by the user). The notebook application may be configured to allow a user to create various projects to complete various tasks. Each project may be configured to act in a manner similar to a folder by which a user can store various information to each project. In some implementations, an individual scratchpad may correspond to or be dedicated to a particular project. In some implementations, the notebook application may be configured to receive the source content as specified by the user. The notebook application may be configured to add, delete, or modify projects according to an input received from a user. Each project may be provided a default name, a name provided by the user, or a name generated by the notebook application (e.g., via one or more machine- learned models) based on the information stored in the project (e.g., based on the source content).

[0065] In some implementations, in response to source content being provided to the notebook application, the notebook application may be configured to automatically generate (e.g., using one or more machine-learned models, one or more generative machine-learned models, etc.), a graphical image (e.g., an emoji, an icon, etc.) or graphical animation which corresponds to or represents the source content. In some implementations, the graphical image or graphical animation may be overlaid on a folder which is provided as a user interface element that, when selected, causes the folder to open and display the contents of the folder to the user. In addition, or alternatively, in some implementations, in response to the source content being provided to the notebook application, the notebook application may be configured to automatically generate (e.g., using one or more machine-learned models, one or more generative machine-learned models, etc.), a textual description (name) which corresponds to or represents the source content. The textual description may be overlaid on the folder which is provided as a user interface element that, when selected, causes the folder to open and display the contents of the folder to the user.

[0066] One or more technical benefits of the disclosure include generating content via one or more machine-learned models, based on particular items of content selected by a user. Current methods for a large language model (LLM) to generate an output require a user tocopy and paste content from one document into a chat box to query the LLM about the content. Switching between multiple windows or applications results in significant amounts of wasted computational time and resources (e.g., processor cycles). In contrast to current methods, a summary regarding user-selected items of content (e.g., source content) can be automatically generated via a notebook application and one or more machine-learned models, in response to a user uploading the items of content. Therefore, a user need not switch between applications or windows, or provide a prompt.

[0067] Another technical benefit of the disclosure includes one or more machine-learned models providing suggested queries based on items of content selected by the user, suggested key topics based on items of content selected by the user, selectable chips based on an output of a response, and the like. A user can select a suggested query and the one or more machine-learned models may be configured to provide a response to the query based on the items of content selected by the user. Providing the suggested query automatically saves computing resources (e.g., networking resources including bandwidth, processor cycles, etc.) by not requiring the user to input the suggested query).

[0068] Another technical benefit of the disclosure includes one or more machine-learned models generating content based on items of content selected by the user and / or based on notes selected by the user. Generation of the content can save time and computing resources by not requiring a user to cut and paste content from multiple sources to generate new content (e.g.. an outline, an essay, a report, etc.) which is based on a plurality of items of content.

[0069] Another technical benefit of the disclosure includes one or more machine-learned models generating a graphical image or animation to display in an overlaid manner on a folder to indicate content which is saved in the folder. The graphical image or animation can improve search capabilities and save computing resources that may otherwise be expended by a user opening and closing folders which do not contain content that the user is actually looking for.

[0070] Another technical benefit of the disclosure includes one or more machine-learned models providing suggested queries based on items of content selected by the user, suggested key topics based on items of content selected by the user, selectable chips based on an output of a response, and the like. A user can select a suggested query and the one or more machine-learned models may be configured to provide a response to the uery based on the items of content selected by the user. Providing the suggested query automatically saves computing resources (e.g.. networking resources including bandwidth, processor cycles, etc.) by not requiring the user to input the suggested query).

[0071] Another technical benefit of the disclosure includes one or more machine-learned models generating a document template based on training documents which can be provided by a user. For example, a user may not be familiar with the typical (accepted) or standard format for certain document types (e.g., a resume, a PRD, a legal opinion, etc.), and may not have the knowledge or expertise to provide an appropriate prompt to generate a document template having the proper format for a particular document type. Therefore, the document extractor application described herein can improve the efficiency of the one or more machine-learned models by generating a document template having the appropriate structure (e.g., style, intent, format, etc.) that can be used to generate an output document based on the proper document template. Accordingly, reduced interactions with the user for generating a document template can save or conserve bandwidth, network resources, computing processing power, etc. Further, reduced inferences for generating the document template by the one or more machine-learned models can also save or conserv e bandwidth, network resources, computing processing power, etc.

[0072] Another technical benefit of the disclosure includes one or more machine-learned models generating an output document based on source documents which can be provided by a user and a document template that can be selected by the user or the computing device. For example, a user may not be familiar with the typical (accepted) or standard format for certain document types (e.g., a resume, a PRD, a legal opinion, etc.), and may not have the knowledge or expertise to provide an appropriate prompt to generate an output document having the proper format for a particular document type. Therefore, the document extractor application described herein can improve the efficiency of the one or more machine-learned models by generating an output document having the appropriate structure (e.g., style, intent, format, etc.) based on a document template that is appropriate for the document type. Accordingly, reduced interactions with the user for generating an output document based on a generated document template via one or more machine-learned models can save or conserve bandwidth, network resources, computing processing power, etc. Further, reduced inferences for generating the output document by the one or more machine-learned models can also save or conserve bandwidth, network resources, computing processing power, etc.

[0073] Referring now to the drawings, FIG. 1 A is an example system according to one or more example embodiments of the disclosure. FIG. 1 A illustrates an example of a system 1000 which includes a computing device 100, an external computing device 200, a server computing system 300, and external content 500, which may be in communication with one another over a network 400. For example, the computing device 100 and the externalcomputing device 200 can include any of a personal computer, a smartphone, a tablet computer, a laptop, a global positioning service device, a smartwatch, and the like. The network 400 may include any type of communications network including a wired or wireless network, or a combination thereof. The network 400 may include a local area network (LAN), wireless local area network (WLAN), wide area network (WAN), personal area network (PAN), virtual private network (VPN), or the like. For example, wireless communication between elements of the example embodiments may be performed via a wireless LAN, Wi-Fi, Bluetooth, ZigBee, Wi-Fi direct (WFD), ultra wideband (UWB), infrared data association (IrDA), Bluetooth low energy (BLE), near field communication (NFC), a radio frequency (RF) signal, and the like. For example, wired communication between elements of the example embodiments may be performed via a pair cable, a coaxial cable, an optical fiber cable, an Ethernet cable, and the like. Communication over the network 400 can use a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e g., VPN, secure HTTP, SSL).

[0074] As will be explained in more detail below, in some implementations the computing device 100 and / or server computing system 300 may form part of an application system which can provide a tool for users to create, manage, or organize information (e.g., documents, imagery, etc ), for example, via one or more machine-learned models.

[0075] In some example embodiments, the server computing system 300 may obtain data from one or more of a content data store 350, a user data store 360, and a machine-learned model data store 370, to implement various operations and aspects of the application system as disclosed herein. The content data store 350, user data store 360, and machine-learned model data store 370 may be integrally provided with the server computing system 300 (e.g., as part of the one or more memory devices 320 of the server computing system 300) or may be separately (e.g., remotely) provided. Further, content data store 350, user data store 360, and machine-learned model data store 370 can be combined as a single data store (database), or may include a plurality of respective data stores. Data stored in one data store (e.g., the content data store 350) may overlap with some data stored in another data store (e.g.. the user data store 360). In some implementations, one data store (e.g., the machine-learned model data store 370) may reference data that is stored in another data store (e.g., the user data store 360).

[0076] In some examples, the content data store 350 can store any kind of information or content. For example, the content data store 350 can include books, product manuals,resumes, legal opinions, academic papers, proprietary data files, patent documents, web pages, emails, forum posts, social media posts, videos, images, geographic information, or any other type or manner of content which may be stored or accessed in digital form (e.g., in a database, memory device, etc.). In some implementations, information may be stored in the content data store 350 by the user selecting certain documents, images, or other content to store in the content data store 350.

[0077] In some examples, the user data store 360 can include information regarding one or more user profiles, including a variety of user data such as user preference data, user demographic data, user calendar data, user social network data, user historical travel data, and the like. For example, the user data store 360 can include, but is not limited to. email data including textual content, images, email-associated calendar information, or contact information; social media data including comments, reviews, check-ins, likes, invitations, contacts, or reservations; calendar application data including dates, times, events, description, or other content; virtual wallet data including purchases, electronic tickets, coupons, or deals; scheduling data; location data; SMS data; or other suitable data associated with a user account. According to one or more examples of the disclosure, the data can be analyzed to determine preferences of the user with respect to generating, managing, and / or organizing content, for example, to automatically generate a summary of a document in a particular manner or style, to automatically provide customized features with respect to content, to provide suggestions, recommendations, and / or questions relating to certain content identified by the user as source content, to automatically generate a document template, to automatically generate a document based on the document template, etc.

[0078] The user data store 360 is provided to illustrate potential data that could be analyzed, in some embodiments, by the computing device 100 and / or server computing system 300 to identify user preferences, to make recommendations, to generate, manage, and / or organize content, etc.. However, such user data may not be collected, used, or analyzed unless the user has consented after being informed of what data is collected and how such data is used. Further, in some embodiments, the user can be provided with a tool (e.g.. in a notebook application or via a user account) to revoke or modify the scope of permissions. In addition, certain information or data can be treated in one or more ways before it is stored or used, so that personally identifiable information is removed or stored in an encrypted fashion. Thus, particular user information stored in the user data store 360 may or may not be accessible to the computing device 100 and / or server computing system 300 based on permissions given by the user, or such data may not be stored in the user data store 360 at all.

[0079] Machine-learned model data store 370 can store machine-learned models which can be retrieved and implemented by the server computing system 300 for generating distilled or fine-tuned machine-learned models (e.g., distilled or fine-tuned generative machine-learned models) that, in some implementations, can also be provided to the computing device 100. Machine-learned model data store 370 can also store distilled or fine-tuned machine-learned models (e.g., distilled or fine-tuned generative machine-learned models) which can be retrieved and implemented by the computing device 100. In some implementations, the computing device 100 can retrieve and implement machine-learned models which are large parameter models that have not been fine-tuned or distilled. The machine-learned models (including large parameter models and distilled or fine-tuned models) stored at the machine- learned model data store 370 can include generative machine-learned models respectively associated with different t pes of content (e.g., different genres or subjects, different kinds of content including imagery, videos, and text, different types of content or documents (e.g., outlines, reports, spreadsheets, etc ), documents having different styles (e.g., casual, opinionated, expert, etc.), documents having different formats (e.g., an outline format, a PRD format, a resume format, etc.), documents having different intentions (e.g.. to inform, to persuade, to compare, to contrast, etc ). The machine-learned models may include large language models (e.g., the Bidirectional Encoder Representations from Transformers (BERT) large language model) and general, multimodal models (e.g., Gemini). The machine-learned models may include generative artificial intelligence (Al) models (e.g.. Bard) which may implement generative adversarial networks (GANs), transformers, variational autoencoders (VAEs), neural radiance fields (NeRFs), and the like.

[0080] External content 500 can be any form of external content including news articles, webpages, video files, audio files, written descriptions, ratings, game content, social media content, photographs, commercial offers, transportation method, weather conditions, sensor data obtained by various sensors, or other suitable external content. The computing device 100, external computing device 200, and server computing system 300 can access external content 500 over network 400. External content 500 can be searched by computing device 100, external computing device 200. and server computing system 300 according to known searching methods and search results can be ranked according to relevance, popularity, or other suitable attributes, including location-specific filtering or promotion.

[0081] Referring now to FIG. IB, example block diagrams of a system 1000’ including a computing device 100 and server computing system 300 according to one or more example embodiments of the disclosure will now be described. Although computing device 100 isrepresented in FIG. IB, features of the computing device 100 described herein are also applicable to the external computing device 200.

[0082] The computing device 100 may include one or more processors 1 10, one or more memory devices 120, an application system 130, a position determination device 140, an input device 150, a display device 160, an output device 170, and a capture device 180. The server computing system 300 may include one or more processors 310, one or more memorydevices 320, and an application system 330.

[0083] For example, the one or more processors 110, 310 can be any suitable processing device that can be included in a computing device 100 or server computing system 300. For example, the one or more processors 110, 310 may include one or more of a processor, processor cores, a controller and an arithmetic logic unit, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an image processor, a microcomputer, a field programmable array, a programmable logic unit, an applicationspecific integrated circuit (ASIC), a microprocessor, a microcontroller, etc., and combinations thereof, including any other device capable of responding to and executing instructions in a defined manner. The one or more processors 110. 310 can be a single processor or a plurality of processors that are operatively connected, for example in parallel.

[0084] The one or more memory devices 120, 320 can include one or more non-transitory computer-readable storage mediums, including a Read Only Memory (ROM), Programmable Read Only Memory (PROM), Erasable Programmable Read Only Memory (EPROM), and flash memory, a USB drive, a volatile memory device including a Random Access Memory (RAM), a hard disk, floppy disks, a blue-ray disk, or optical media such as CD ROM discs and DVDs, and combinations thereof. However, examples of the one or more memorydevices 120, 320 are not limited to the above description, and the one or more memory devices 120, 320 may be realized by other various devices and structures as would be understood by those skilled in the art.

[0085] For example, the one or more memory- devices 120 can also include data 122 and instructions 124 that can be retrieved, manipulated, created, or stored by the one or more processors 110. In some example embodiments, such data can be accessed and used as input to implement notebook application 132, and to execute the instructions to perform operations including: providing a user interface including a first portion and a second portion, wherein the first portion includes a textual summary generated via one or more machine-learned models based on a plurality of documents selected by a user and the second portion includes a plurality of user interface elements to perform an operation with respect to the textualsummary, as described according to examples of the disclosure. In some example embodiments, such data can be accessed and used as input to implement document extractor application 136, and to execute the instructions to perform operations including: receiving a plurality of source documents, receiving an input associated with a request to generate an output document based on the plurality' of source documents and a particular document template, and generating, via one or more machine-learned models, the output document having the particular document template, as described according to examples of the disclosure.

[0086] For example, the one or more memory' devices 320 can also include data 322 and instructions 324 that can be retrieved, manipulated, created, or stored by the one or more processors 310. In some example embodiments, such data can be accessed and used as input to implement notebook application 332, and to execute the instructions to perform operations including: providing a user interface including a first portion and a second portion, wherein the first portion includes a textual summary generated via one or more machine-learned models based on a plurality of documents selected by a user and the second portion includes a plurality of user interface elements to perform an operation with respect to the textual summary, as described according to examples of the disclosure. In some example embodiments, such data can be accessed and used as input to implement document extractor application 336, and to execute the instructions to perform operations including: receiving a plurality of source documents, receiving an input associated with a request to generate an output document based on the plurality of source documents and a particular document template, and generating, via one or more machine-learned models, the output document having the particular document template, as described according to examples of the disclosure.

[0087] In some example embodiments, the computing device 100 includes an application system 130. For example, the application system 130 may include the notebook application 132, a document application 134 (e.g., a word processing application, a spreadsheet application, a presentation application, an imagery’ application, etc.), and a document extractor application 136. The application system 130 can include various other applications including text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, map applications, social media applications, navigation applications, etc.

[0088] According to examples of the disclosure, the notebook application 132 may be executed by the computing device 100 to provide a user of the computing device 100 a wayto organize, manage, create, and interact with content, particularly with content that is curated or selected by the user. In some implementations, the notebook application 132 may be part of document application 134, or may be a standalone application. The notebook application 132 may be configured to be dynamically interactive according to various user inputs. Example implementations of the notebook application 132 are described herein, however the disclosure is not limited to these examples as various modifications may be made to the embodiments described herein.

[0089] In some examples, one or more aspects of the notebook application 132 may be implemented by the notebook application 332 of the server computing system 300 which may be remotely located, to organize, manage, create, and interact with content, in response to receiving an input from a user. In some examples, one or more aspects of the notebook application 332 may be implemented by the notebook application 132 of the computing device 100, to organize, manage, create, and interact with content, in response to receiving an input from a user.

[0090] According to examples of the disclosure, the document application 134 may be executed by the computing device 100 to provide a user of the computing device 100 a way to organize, manage, create, and interact with content, particularly with content that is curated or selected by the user. The document application 134 can be any kind of application that pertains to documents (e.g., in a textual or visual format), and can include word processing applications, spreadsheet applications, presentation applications, visual applications, portable document format file applications, etc. In some implementations, the notebook application 132, document application 134, and document extractor application 136 may interact with each other. For example, content from a document that is created via the document application 334 may be uploaded or stored for use with notebook application 132. In some implementations, the notebook application 132 may be configured to generate a document (e.g., a report, an outline, a presentation, a spreadsheet) which can be compatible with (opened by or exported to) the document application 134. In some implementations, the notebook application 132 may be configured to generate the document (e.g., a report, an outline, a presentation, a spreadsheet) such that the document has a particular style, format, and / or intent, based on a document template that is generated via the document extractor application 136.

[0091] In some examples, the document application 134 can be a dedicated application specifically designed to provide a particular service. In other examples, the documentapplication 134 can be a general application (e.g., a web browser) and can provide access to a variety of different services via the network 400.

[0092] According to examples of the disclosure, the document extractor application 136 may be executed by the computing device 100 to provide a user of the computing device 100 a document template for organizing, managing, creating, and interacting with content, particularly with content that is curated or selected by the user. In some implementations, the document extractor application 136 may be part of document application 134 and / or notebook application 132, or may be a standalone application. The document extractor application 136 may be configured to be dynamically interactive according to various user inputs. Example implementations of the document extractor application 136 are described herein, however the disclosure is not limited to these examples as various modifications may be made to the embodiments described herein.

[0093] In some examples, one or more aspects of the document extractor application 136 may be implemented by the document extractor application 336 of the sen' er computing system 300 which may be remotely located, to provide a document template for organizing, managing, creating, and interacting with content, in response to receiving an input from a user. In some examples, one or more aspects of the document extractor application 336 may be implemented by the document extractor application 136 of the computing device 100, to provide a document template for organizing, managing, creating, and interacting with content, in response to receiving an input from a user.

[0094] In some example embodiments, the computing device 100 includes a position determination device 140. Position determination device 140 can determine a current geographic location of the computing device 100 and communicate such geographic location to server computing system 300 over network 400. The position determination device 140 can be any device or circuitry for analyzing the position of the computing device 100. For example, the position determination device 140 can determine actual or relative position by using a satellite navigation positioning system (e.g. a GPS system, a Galileo positioning system, the GLObal Navigation satellite system (GLONASS), the BeiDou Satellite Navigation and Positioning system), an inertial navigation system, a dead reckoning system, based on an IP address, by using triangulation and / or proximity to cellular towers or WiFi hotspots, and / or other suitable techniques for determining a position of the computing device 100.

[0095] The computing device 100 may include an input device 150 configured to receive an input from a user and may include, for example, one or more of a keyboard (e.g., a physicalkeyboard, virtual keyboard, etc.), a mouse, a joystick, a button, a switch, an electronic pen or stylus, a gesture recognition sensor (e.g., to recognize gestures of a user including movements of a body part), an input sound device or speech recognition sensor (e.g., a microphone to receive a voice input such as a voice command or a voice query), a track ball, a remote controller, a portable (e.g., a cellular or smart) phone, a tablet PC, a pedal or footswitch, a virtual -reality device, and so on. The input device 150 may also be embodied by a touch- sensitive display having a touchscreen capability, for example. For example, the input device 150 may be configured to receive an input from a user associated with the input device 150 for selecting content that is to be organized or managed, for selecting queries or actions with respect to content that is curated or selected by the user, for uploading a plurality of training documents for generating a document template that can be applied with respect to a plurality of source documents, for selecting the plurality of source documents for generating an output document based on the document template, etc.

[0096] The computing device 100 may include a display device 160 which displays information viewable by the user (e.g., a user interface screen). For example, the displaydevice 160 may be a non-touch sensitive display or a touch-sensitive display. The display device 160 may include a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emiting diode (OLED) display, active matrix organic light emiting diode (AMOLED), flexible display, 3D display, a plasma display panel (PDP), a cathode ray tube (CRT) display, and the like, for example. However, the disclosure is not limited to these example displays and may include other types of displays. The display device 160 can be used by the application system 130 provided at the computing device 100 to display information to a user relating to an input (e.g., information relating to a document, to a note, to a project, to a document template, to a user interface screen having user interface elements which are selectable by the user, etc.).

[0097] The computing device 100 may include an output device 170 to provide an output to the user and may include, for example, one or more of an audio device (e.g., one or more speakers), a haptic device to provide haptic feedback to a user (e.g., a vibration device), a light source (e.g., one or more light sources such as LEDs which provide visual feedback to a user), a thermal feedback system, and the like.

[0098] The computing device 100 may include a capture device 180 that is capable of capturing media content, according to various examples of the disclosure. For example, the capture device 180 can include an image capturer 182 (e.g.. a camera) which is configured to capture images (e.g., photos, video, and the like). For example, the capture device 180 caninclude a sound capturer 184 (e.g., a microphone) which is configured to capture sound or audio (e.g.. an audio recording). The media content captured by the capture device 180 may be transmitted to one or more of the server computing system 300, content data store 350, user data store 360, and machine-learned model data store 370, for example, via network 400. For example, in some implementations, media content which is captured by the capture device 180 may be selected as source content by a user for use in creating a note with respect to a project. The media content can be provided as an input to one or more machine-learned models to generate a note, for example.

[0099] In accordance with example embodiments of the disclosure, the server computing system 300 can include one or more processors 310 and one or more memory devices 320 as described herein. The server computing system 300 may also include an application system 330 which is similar to the application system 130 described herein.

[0100] For example, the application system 330 may include a notebook application 332 which performs functions similar to those discussed herein with respect to notebook application 132, a document application 334 which includes applications similar to those discussed above with respect to document application 134. and a document extractor application 336 which performs functions similar to those discussed herein with respect to document extractor application 136. In some implementations, one or more machine-learned models (e.g., generative machine-learned models, large language models, etc.) associated with the application system 330 may be configured to organize, manage, create, and interact with content based on source content that is curated or selected by a user. For example, one or more machine-learned models (e.g., generative machine-learned models, large language models, etc.) associated with the application system 330 may be configured to perform a first action (e.g., generate a summary or document guide with respect to source content selected by a user), while the computing device 100 may be configured to perform a second action (e.g., generate suggested actions, generate an outline or study guide based on a plurality of notes saved to a scratchpad). For example, one or more machine-learned models (e.g., generative machine-learned models, large language models, etc.) associated with the application system 130 may be configured to perform a first action (e.g., upload source content selected by a user), while the server computing system 300 may be configured to perform a second action (e.g., generate a document template based on the uploaded source content). For example, a particular action to be performed by the application system 330 may vary according to a network status (e.g.. an available bandwidth, a channel utilization status. a latency status, a throughput rate, etc.). In some implementations, one or more machine-learned models associated with the application system 330 may be configured to process a user input to generate information (e.g., semantic information) which can then be provided as an input to one or more other machine-learned models (e.g., generative machine-learned models, large language models, etc.) associated with the application system 330, to generate the content to be utilized with respect to a project for the notebook application 132 and / or notebook application 332.

[0101] Examples of the disclosure are also directed to computer implemented methods for providing a user interface for organizing, managing, and creating content by implementing one or more machine-learned models with respect to source content selected by a user. FIG. 2 illustrates a flow diagram of an example, non-limiting computer-implemented method, according to one or more example embodiments of the disclosure. FIG. 3 illustrates a block diagram of a notebook application, according to one or more example embodiments of the disclosure.

[0102] The flow diagram of FIG. 2 illustrates a method 2000 for providing a user interface for organizing, managing, and creating content by implementing one or more machine-learned models with respect to source content selected by a user. Although shown in a particular sequence or order, unless otherwise specified, the order of the processes can be modified. Thus, the illustrated embodiments should be understood only as examples, and the illustrated processes can be performed in a different order, and some processes can be performed in parallel. Additionally, one or more processes can be omitted in various embodiments. Thus, not all processes are required in every embodiment. Other process flows are possible.

[0103] Referring to FIG. 2, at operation 2100 the method 2000 includes a computing device receiving an input from a user relating to the selection of source content. As described herein, the computing device may be embodied as computing device 100, server computing system 300, or combinations thereof. For example, the input may be provided by the user via input device 150. For example, the input may be provided by selecting particular files or documents which are uploaded to the computing device for use by the notebook application 132. In some implementations, the content can be uploaded from a local memory’, from another application (e.g.. a portable document file application), from copied text, or from a w ebsite. The selected files, text, documents, etc. may’ be referred to as source content. In some implementations, the source content may be a subset of a larger corpus of content. The input may be provided or input to notebook application 132 or notebook application 332, for example.

[0104] In some implementations, a response to the input selecting the source content may be processed at computing device 100 without involving the sen- er computing system 300. In some implementations, the input selecting the source content may be transmitted from computing device 100 to server computing system 300 and at least part of the response to the input may be processed by the server computing system 300. For example, the input relating to the selection of the source content may be provided at the computing device 100 and the server computing system 300 may be configured to perform an operation in response to receiving an indication of the input.

[0105] At operation 2200, the computing device may be configured to implement one or more machine-learned models with respect to the selected source content to generate a document guide. In some implementations, the document guide (source guide) generated by the one or more machine-learned models may include a summary of the source content and key topics relating to the source content. In some implementations, the document guide may further include one or more suggested queries (e.g., questions) that may be provided in the form of a selectable user interface element.

[0106] For example, the computing device can obtain information indicating that the user has selected source content. The computing device can process the source content with one or more machine-learned models (e.g., one or more large language models) to obtain a language output. The computing device can then use the one or more machine-learned models (e.g.. one or more large language models) to generate a summarization output. In particular, a machine-learned large language model can be trained to process a variety of outputs to generate a language output. For example, the machine-learned large language model can process an embedding generated by a machine-learned embedding generation model, portions of the source content identified using the embedding generation model, language outputs generated using the machine-learned large language model or some other model, etc.

[0107] At operation 2300, the computing device may be configured to receive an input to perform an action with respect to the document guide. At operation 2400 the computing device may be configured to perform the action in response to receiving the input. For example, the input may be the selection of a suggested query and the action may include providing an answer to the question by implementing the one or more machine-learned models with respect to the source content. For example, the input may be a text input asking a question and the action may include providing an answer to the question by implementing the one or more machine-learned models with respect to the source content. For example, the input may be a selection of a portion of the summary and the action may include providing anoutput indicating particular sources from among the source content which were relied upon for generating the text associated with the selection of the portion of the summary.

[0108] Referring to FIG. 3, notebook application 3100 (which may correspond to notebook application 132 and / or notebook application 332) may include a conditioning parameters generator 3110, one or more sequence processing models 3120, one or more large language models 3130, and one or more generative machine-learned models 3140. The notebook application 3100 may receive an input 3200 from a user as discussed above with respect to operation 2100 and operation 2300 of FIG. 2. Conditioning parameters generator 3110 may be configured to generate conditioning parameters based at least in part on the input, wherein the conditioning parameters provide values for one or more conditions associated with content to be generated which relates at least in part to the input 3200 and source content 3400 selected by the user.

[0109] For example, source content 3400 can include any kind of document (e.g., in digital form) and may include books, product manuals, legal opinions, academic papers, proprietary data files, patent documents, web pages, emails, forum posts, social media posts, videos, images, geographic information, or any other type or manner of content which may be stored or accessed in digital form (e g., in a database, memory device, etc.). In some implementations, source content 3400 may be stored in the content data store 350 by the user selecting certain documents, images, or other content to store in the content data store 350. In some implementations, source content 3400 may be stored at the computing device 100 or server computing system 300.

[0110] To generate the conditioning parameters, the conditioning parameters generator 3110 may be configured to retrieve values for the one or more conditions associated with the input. For example, to generate the conditioning parameters, the conditioning parameters generator 3110 may be configured to extract the values for the one or more conditions from the input. The input may include information indicative of the user's intent or requirements. In some implementations, the conditioning parameters generator 3110 (or the one or more sequence processing models 3120 or the one or more large language models 3130) may be configured to extract information from the input 3200 to identify values for the one or more conditions, and the conditioning parameters generator 3110 may be configured to generate the conditioning parameters based on the extracted values. For example, the input itself may identify' a color to be used for headings in a generated document (e.g., “blue font for the title”) or an attribute or feature (e.g.. “circle bullet points”) that can be used to generate the conditioning parameters for generating a document related to the source content.

[0111] To generate the conditioning parameters, the conditioning parameters generator 3110 may be configured to infer the values for the one or more conditions from the input. The input may include information indicative of the user's intent or requirements. In some implementations, the conditioning parameters generator 3110 (or the one or more sequence processing models 3120 or the one or more large language models 3130) may be configured to infer information from the input 3200 to identify values for the one or more conditions, and the conditioning parameters generator 3110 may be configured to generate the conditioning parameters based on the inferred values. For example, the input may include a reference to a length (“short,” “long,” etc.) of the summan' to be generated or of another document to be generated based on the source content, and the conditioning parameters generator 3110 (or the one or more sequence processing models 3120 or the one or more large language models 3130) may be configured to infer a value based on the input. For example, an input requesting the notebook application 3100 to generate a “short” essay may infer a value of about 500 words while a “long” essay may be associated with a value of about 2000 words. For example, the notebook application 3100 may be configured to ascertain an inferred value based on information via external content 3300.

[0112] In some implementations, the conditioning parameters generator 3110 may be configured to infer the values for the one or more conditions from the input by providing the input to one or more sequence processing models 3120, w herein the one or more sequence processing models 3120 are configured to output the values for the one or more conditions in response to or based on the query. The one or more sequence processing models 3120 may include one or more machine-learned models which are configured to process and analyze sequential data and to handle data that occurs in a specific order or sequence, including time series data, natural language text, or any other data with a temporal or sequential structure.

[0113] The one or more sequence processing models 3120 may receive an input including text and tokenize the input by breaking down the sequence of text into small units (tokens) to provide a structured representation of the input sequence. The one or more sequence processing models 3120 may represent the tokens as vectors in a continuous vector space by mapping each token to a high-dimensional vector, where the relationships betw een tokens (words) are reflected in the geometric relationships between their corresponding vector. For example, the one or more sequence processing models 3120 may receive an input including the text “How did the Cold War end?” and tokenize the input by breaking down the sequence of text into small units (tokens) (e.g., “How.” “Cold War.” and “end”), thereby providing a structured representation of the input sequence. In a word embedding, semantically similarwords are closer together in the vector space. For example, the vectors for "war" and "battle" might be close to each other because of their semantic relationship, while the vectors for ■‘war” and ‘'peace” may be far apart compared to the vectors for “war” and “battle”.

[0114] The one or more large language models 3130 can be, or otherwise include, a model that has been trained on a large corpus of language training data in a manner that provides the one or more large language models 3130 with the capability to perform multiple language tasks. For example, the one or more large language models 3130 can be trained to perform summarization tasks, conversational tasks, simplification tasks, oppositional viewpoint tasks, etc. In particular, the one or more large language models 3130 can be trained to process a variety of outputs to generate a language output. For example, the one or more large language models 3130 can process an embedding generated by a machine-learned embedding generation model, portions of source content (e.g., document chunk(s)) identified using an embedding generation model, language outputs generated using the one or more large language models 3130 or some other model, etc.

[0115] The one or more generative machine-learned models 3140 may include a deep neural network or a generative adversarial network (GAN), variational autoencoders, stable diffusion machine-learned models, visual transformers, neural radiance fields (NeRFs), etc., to generate content (e.g., a summary', response to a query', etc.) with values for conditions associated with one or more features. For example, the computing device may include a database (e.g., machine-learned model data store 370) which is configured to store a plurality of generative machine-learned models respectively associated with a plurality of different ty pes of content (e.g., different genres or subjects, different kinds of content including imagery', videos, and text, different styles of content including outlines, reports, spreadsheets, etc.). In some implementations, the computing device may be configured to retrieve, from among the one or more generative machine-learned models 3140, a generative machine- learned model associated with a particular type of content relating to the input.

[0116] In some implementations, the one or more generative machine-learned models 3140 may be trained on a large dataset of content (e.g., a large corpus of language training data) with corresponding information about the conditions associated with the content. During training, the one or more generative machine-learned models 3140 leam relationships between elements in an output (e.g., content) and conditions that influence them. This may involve the computing device adjusting each generative machine-learned model’s internal parameters to generate realistic or accurate content (e.g., grammatically correct content, coherent content, etc.) based on the training data. The one or more generative machine-learned models 3140 may be trained on one or more training datasets including a plurality of reference images of the location. The one or more training datasets may include values for the one or more conditions.

[0117] In some implementations, the one or more generative machine-learned models 3140 are configured to generate the document guide 3500 in response to receiving the selection of source content 3400 and / or to generate responsive content 3600 which corresponds to content that is generated in response to the input to perform an action with respect to the document guide, etc., based on the conditioning parameters (and corresponding values for the one or more conditions) to make decisions for generating content.

[0118] In some implementations, the server computing system 300 may provide (transmit) content or a portion of the generated content to computing device 100 or the server computing system 300 may provide access to the generated content to the computing device 100. For example, the document guide 3500 may be generated at the server computing system 300 and stored at one or more computing devices (e g., one or more of computing device 100, external computing device 200, server computing system 300, external content 500, content data store 350, user data store 360. etc.).

[0119] In some implementations, after a document guide is generated and / or after an action is performed with respect to the document guide, the user can provide feedback or a further input relating to the content which is generated based on the source content provided and / or a query provided via the user, and one or more of the operations 2100 through 2400 can be repeated.

[0120] Examples of the disclosure are also directed to user-facing aspects by which a user can manage content, organize content, create content, etc., via a notebook application which is configured to implement one or more machine-learned models with respect to source content selected by the user. For example, FIGS. 4A through 4H illustrate examples of actions which can be implemented for a project in which a document guide is generated via one or more machine-learned models based on source content selected by a user, according to one or more example embodiments of the disclosure.

[0121] For example. FIG. 4A illustrates a first user interface screen (e.g., a startup user interface screen, a startup graphical user interface, etc.) of a notebook application, according to one or more example embodiments of the disclosure.

[0122] In FIG. 4A, first user interface screen 4100 depicts a user interface (e.g., a launch screen) which provides information about the notebook application 3100. In particular, notebook application 3100 is configured to present for display the first user interface screen4100 which includes various information 4110 regarding features which are available in the notebook application 3100.

[0123] As illustrated in FIG. 4B, the notebook application 3100 is further configured to present for display a second user interface screen 4200 which includes a first user interface element 4210. For example, the first user interface element 4210 is associated with enabling a user to create a new notebook (a new project) by which a user can manage content, organize content, create content, etc., based on source content which the user can select or curate.

[0124] As illustrated in FIG. 4C, the notebook application 3100 is further configured to present for display a third user interface screen 4300 in response to a user providing an input to create anew notebook (e.g., via the selection of the first user interface element 4210). The third user interface screen 4300 includes a first portion 4310 having a plurality of selectable user interface elements that correspond to locations where the source content can be uploaded from. For example, first user interface element 4312 corresponds to a storage space which may be associated with a local computing device or a remote server system (e.g., a cloud server), or another storage device (e.g., a portable storage device). For example, second user interface element 4314 corresponds to a portable document format file, third user interface element 4316 corresponds to copied text, and fourth user interface element 4318 corresponds to content which can be uploaded from a particular website or URL.

[0125] As illustrated in FIG. 4D, the notebook application 3100 is further configured to present for display a fourth user interface screen 4400 in response to the selection of one of the plurality of selectable user interface elements that correspond to locations where the source content can be uploaded from, described with respect to FIG. 4C. The fourth user interface screen 4400 includes a first portion 4410 having a plurality of selectable items of source content (e.g., a plurality’ of documents, images, videos, etc.). For example, first user interface element 4412 corresponds to a first selected document, second user interface element 4414 corresponds to a second selected document (e.g., a portable document format file), and third user interface element 4416 corresponds to a third selected document. FIG. 4D illustrates that the user can curate or select particular items of source content which can be used for creating a notebook or project and which can be relied upon by one or more machine-learned models as input data for organizing content, managing content, creating content, etc.

[0126] As illustrated in FIG. 4E, the notebook application 3100 is further configured to present for display a fifth user interface screen 4500 in response to the selection of one or more items of content from the plurality of items of source content, described with respect to FIG. 4D.The fifth user interface screen 4500 includes a first portion 4510, a second portion 4520, a third portion 4530. and a fourth portion 4540. Each portion of the fifth user interface screen 4500 may correspond to a section or panel of the fourth user interface screen 4400 and can be associated with a different functionality.

[0127] For example, the first portion 4510 corresponds to a document guide (also referred to as a source guide) which includes a summary’ section 4512 and a key topics section 4514. The notebook application 3100 may be configured to generate the content (e.g.. a textual description) associated with the summary section 4512 by implementing one or more machine-learned models as described herein with respect to FIG. 3, based on the selected source content (e.g., as described with respect to FIG. 4D). For example, the summary section 4512 may provide a brief summary associated with one or more of the items of content which comprise the selected source content. Likewise, the notebook application 3100 may be configured to generate the content (e.g., a textual description) associated with the key topics section 4514 by implementing one or more machine-learned models as described herein with respect to FIG. 3, based on the selected source content (e.g., as described with respect to FIG. 4D). For example, the key topics section 4514 may include one or more user interface elements which identify themes or important topics associated with one or more of the items of content which comprise the selected source content. Further, the notebook application 3100 may be configured to generate an output in response to a selection of one of the user interface elements in the key topics section 4514. The output may be a text summary or text explanation regarding the key topic corresponding to the selected user interface element, for example. The output may be provided in a separate user interface screen or provided in another portion of the fifth user interface screen which the notebook application 3100 is configured to generate in response to the selection of one of the user interface elements in the key topics section 4514.

[0128] For example, the second portion 4520 corresponds to a source content section (e.g., a context window) which includes information 4522 from at least a portion of an item of content from the source content. The notebook application 3100 may be configured to reproduce at least a portion of an item of content from the source content in the second portion 4520. In some implementations, the content in the source content section may correspond to a portion of an item of content which was relied upon for generating the summary section 4512.

[0129] For example, the third portion 4530 corresponds to a notes section (e.g., a scratchpad) which can include one or more notes that may be generated via various methods as describedherein (e.g., automatically generated by the notebook application 3100, manually entered by a user, automatically generated by the notebook application 3100 in response to the selection of a user interface element which corresponds to an action to be performed, etc.).

[0130] For example, the fourth portion 4540 corresponds to a query section which can include one or more user interface elements for submitting or providing a query' to the notebook application 3100 with respect to the source content. For example, the fourth portion 4540 includes a plurality of user interface elements 4542 which correspond to suggested questions or actions that are related to the source content. For example, the notebook application 3100 may be configured to generate the suggested questions or actions based on information included in the source content. For example, the notebook application 3100 may be configured to generate the suggested questions or actions based additionally on dialogue history (e.g., prior questions or queries), user data (e.g., preferences of the user, user attributes, etc.), and other contextual information. The fourth portion 4540 may further include a text entry box 4544 by which a user can provide an input (e.g., via a keyboard, via a voice input, etc.) to query the notebook application 3100. The fourth portion 4540 may further include a user interface element 4546 which indicates the number of items of content which comprise the source content. For example, in FIG. 4E, user interface element 4546 indicates three sources were relied upon by the notebook application 3100 to generate the summary section 4512.

[0131] Referring to FIG. 4F, an example user interface screen illustrates an input question and output response relating to the source content. For example, in FIG. 4F the notebook application 3100 is further configured to present for display a sixth user interface screen 4600 in response to receiving a query (e.g., a text query input via the text entry box 4544 of FIG. 4E). For example, sixth user interface screen 4600 includes a first portion 4610 which corresponds to a dialogue section, a second portion 4620 which corresponds to a sources section, and a third portion 430 which corresponds to a notes section (e.g., a scratchpad).

[0132] For example, the first portion 4610 includes a prompt area 4612 that corresponds to the text query and a response area 4614 that corresponds to the response to the text query’. In some implementations, the notebook application 3100 is configured to generate the response by implementing one or more machine-learned models in response to receiving the text query as an input and with reference to the source content 3400. For example, if a user inputs a question (e g., ‘‘How did the Cold War affect American foreign policy?) via the text entry' box 4544 as described with respect to FIG. 4E, the notebook application 3100 may be configured to provide the sixth user interface screen 4600 and to generate a response as indicated in theresponse area 4614. As indicated in the response area 4614, the number of references (items of source content) relied upon by the one or more machine-learned models to generate the response may be indicated by a first user interface element 4616. In the example of FIG. 4F, three references were used to generate the response. The response area 4614 further includes a selectable second user interface element 4618 that, when selected, causes the response to be saved as a note to the third portion 4630 which corresponds to the notes section (e.g., a scratchpad) which can include one or more notes that may be generated via various methods as described herein (e g., automatically generated by the notebook application 3100 in response to the selection of second user interface element 4618, etc.).

[0133] In some implementations, one or more portions of the response area may include information which is selectable that, when selected, can cause additional information to be displayed relating to the selected information. For example, in FIG. 4F the text “policy of containment” may be highlighted, bolded, underlined, or be displayed in some visually distinct manner to indicate that the text is selectable (e.g., a clickable chip) and additional information relating to the text is available. The notebook application 3100 may be configured to provide the additional information (e.g., by implementing one or more machine-learned models based on the selected source content) to provide additional information relating to the text, in response to the selection of the text.

[0134] The second portion 4620 may correspond to a source section and include the items of content 4622 which comprise the source content. In some implementations, the items of content 4622 may correspond to items of content which are relied upon by the one or more machine-learned models for generating the response. In some implementations, the notebook application 3100 may be configured to dynamically modify or re-generate a response in the response area 4614, in response to receiving an additional item of content to be added as source content via the user interface element 4624. In addition, or alternatively, in some implementations, the notebook application 3100 may be configured to dynamically modify or re-generate a response in the response area 4614, in response to receiving a deselection of an item of content from the list of items of content in the second portion 4620 via the user interface element 4626 (e.g.. by unchecking the checkbox for one or more of the items of content in the second portion 4620).

[0135] Referring to FIG. 4G, an example user interface screen includes an example notes section (scratchpad) for a project, according to examples of the disclosure. For example, in FIG. 4G the notebook application 3100 is further configured to present for display a seventh user interface screen 4700 in response to receiving a selection of the second user interfaceelement 4618 (e.g., as shown in FIG. 4F) that, when selected, causes the response to be saved as a note 4712 to the third portion 4710 which corresponds to the notes section (e.g., a scratchpad) which can include one or more notes that may be generated via various methods as described herein (e.g., automatically generated by the notebook application 3100 in response to the selection of second user interface element 4618, etc.). In FIG. 4G, user interface element 4714 indicates the number of items of content the one or more machine- learned models relied upon to generate the response for note 4712. Further, user interface element 4714 may be configured to be selectable such that in response to user interface element 4714 being selected, a list of the items of content (citations) from the source content used for generating the response can be provided for display.

[0136] Referring to FIG. 4H, an example user interface screen includes a notes section (scratchpad) for a project, according to examples of the disclosure. For example, in FIG. 4H the notebook application 3100 is further configured to present for display an eighth user interface screen 4800 in response to receiving a selection of an item 4816 of content from a list 4814 of items of content (citations) from the source content used by the one or more machine-learned models for generating the response saved in the note 4812 which is provided for display in the first portion 4810. The eighth user interface screen 4800 further includes a second portion 4820 which corresponds to a sources section. In FIG. 4H, the notebook application 3100 is configured to provide for display in the second portion 4820 information 4822 relating to the selected item 4816, in response to receiving the selection of the item 4816 of content from the list 4814 of items of content (citations) from the source content used by the one or more machine-learned models for generating the response saved in the note 4812.

[0137] In some implementations the information 4822 may include information from the item 4816 of content that was used to generate the response. For example, the notebook application 3100 may be configured to reference metadata associated with the response to refer back to the information 4822. The metadata may indicate a location of information from an item of content used to generate the response. Further, the information 4822 may correspond to or include a particular passage that was relied upon from the item of content for generating the response. For example, the notebook application 3100 may be configured to cause the particular passage to be displayed in the second portion 4820 in a visually distinctive manner (e.g., in a highlighted manner, a bold manner, an enlarged font size, an underlined manner, an italicized manner, etc.). For example, the notebook application 3100 may be configured to cause additional passages which appear before and / or after theparticular passage to be displayed in the second portion 4820. This additional information may provide further context for the user regarding the information that was relied upon for generating the response. For example, the notebook application 3100 may be configured to mark particular items of content relied upon for generating the response in the note 4812 as well as mark particular passages from the particular items of content relied upon for generating the response in the note 4812. Therefore, a user can easily and visually discern where support for a response can be found in an item of content.

[0138] In some implementations, the information 4822 from the selected item 4816 of content that was used to generate the response may be truncated or shown in its entirety. For example, when the information 4822 is less than a threshold value, the entire text from the selected item 4816 of content can be shown in the second portion 4820 and can be used by the one or more machine-learned models for generating a response (e.g., to a text query). For example, when the information 4822 is more than the threshold value, the notebook application 3100 may be configured to implement a semantic retrieval method to determine particular passages from the entirety of the selected item 4816 of content which are relevant to a user query (e.g.. a text query ). In this example, the relevant passages (rather than the entirety' of the information from the item of content) is relied upon by the one or more machine-learned models for generating a response to the user query' (e.g., the text query').

[0139] Examples of the disclosure are directed to further user-facing aspects by which a user can manage content, organize content, create content, etc., via a notebook application which is configured to implement one or more machine-learned models with respect to source content selected by the user. For example, FIGS. 5A through 5B illustrate examples of actions which can be implemented for a project in which a note is generated via one or more machine-learned models based on source content selected by a user, according to one or more example embodiments of the disclosure.

[0140] For example, FIG. 5A illustrates a first user interface screen of a notebook application, according to one or more example embodiments of the disclosure. For example, in FIG. 5A the first user interface screen 5100 includes a first portion 5110, a second portion 5120, and a third portion 5130. First portion 5110 corresponds to a notes section (e.g., a scratchpad) which can include one or more notes 51 12 that may be generated via various methods as described herein (e.g., automatically generated by the notebook application 3100, manually entered by a user, automatically generated by the notebook application 3100 in response to the selection of a user interface element which corresponds to an action to be performed, etc.).

[0141] Second portion 5120 corresponds to a source content section (e.g., a source guide or context window) which can include one or more sources 5122 (e.g., items of content which comprises the source content 3400 relied upon by the one or more machine-learned models for generating the information included in the one or more notes 5112). In some implementations, the notebook application 3100 may be configured to generate a note which is saved to the first portion 5110 as a note based on a selection of at least a portion of the information from an item of content which is provided in the second portion 5120. For example, FIG. 5 A illustrates selected text 5124 (e.g., highlighted text) that has been selected by a user.

[0142] For example, the third portion 5130 corresponds to a query section which can include one or more user interface elements for submitting or providing a query to the notebook application 3100 with respect to the source content. For example, the third portion 5130 includes a plurality of user interface elements 5132 which correspond to suggested questions or actions that are related to the source content. For example, the notebook application 3100 may be configured to generate the suggested questions or actions based on information included in the source content and / or based on the information displayed in the second portion 5120. The third portion 5130 may further include atext entry box 5134 by which a user can provide an input (e.g., via a keyboard, via a voice input, etc.) to query7the notebook application 3100. The third portion 5130 may further include a user interface element 5136 which indicates the number of items of content which comprise the source content. For example, in FIG. 5 A, user interface element 5136 indicates three sources were relied upon by the notebook application 3100 to generate the one or more notes 5112.

[0143] In some implementations, the plurality of user interface elements 5132 may be configured to dynamically change based on actions with respect to the first user interface screen 5100. For example, the notebook application 3100 may be configured to dynamically change, modify, delete, or add user interface elements in the third portion 5130 based on an action with respect to the source content (e.g., with respect to items of content provided for display in the second portion 5120). In FIG. 5A, the notebook application 3100 may be configured to dynamically change user interface elements in the third portion 5130 based on (in response to) the selection of text from one or more sources 5122 (e.g., the selected text 5124). For example, as indicated in FIG. 5A the actions may include summarizing the selected text to a note, adding a quote to a note, requesting additional information regarding the selected text 5124, or suggest related ideas. For example, the notebook application 3100 may be configured to generate a note summarizing the selected text in response to receiving aselection of user interface element 5132a which corresponds to the action of summarizing the selected text to a note. For example, the notebook application 3100 may be configured to add content to an existing note corresponding to the selected text in response to receiving a selection of user interface element 5132b which corresponds to the action of adding a quote to a note.

[0144] For example, FIG. 5B illustrates a second user interface screen of a notebook application, according to one or more example embodiments of the disclosure. For example, in FIG. 5B second user interface screen 5200 includes a first portion 5210, a second portion 5220, and a third portion 5230, each of which may correspond to the first portion 5110, second portion 5120. and third portion 5130 of FIG. 5A.

[0145] As described with respect to FIG. 5 A. the notebook application 3100 may be configured to generate a note summarizing the selected text in response to receiving a selection of user interface element 5132a which corresponds to the action of summarizing the selected text to a note. FIG. 5B illustrates the generated note 5214 which has been saved to the first portion 5210 which includes one or more notes 5212. Further, in some implementations after the generated note 5214 is saved to the first portion 5210, the plurality of user interface elements 5132 from FIG. 5 A may be configured to dynamically change back to a previous state to the plurality of user interface elements 5232 shown in FIG. 5B.

[0146] Examples of the disclosure are directed to further user-facing aspects by which a user can manage content, organize content, create content, etc., via a notebook application which is configured to implement one or more machine-learned models with respect to source content selected by the user. For example, FIGS. 6A through 6B illustrate examples of actions which can be implemented for a project in which a note is generated via one or more machine-learned models based on source content selected by a user, according to one or more example embodiments of the disclosure.

[0147] For example, FIG. 6A illustrates a portion of a first user interface screen of a notebook application, according to one or more example embodiments of the disclosure. For example, in FIG. 6A a first portion 6110 and a second portion 6120 of a user interface screen are shown. First portion 6110 corresponds to a notes section (e.g., a scratchpad) which can include a plurality of notes that may have been generated via various methods as described herein (e.g., automatically generated by the notebook application 3100, manually entered by a user, automatically generated by the notebook application 3100 in response to the selection of a user interface element which corresponds to an action to be performed, etc.). For example,the first portion 6110 may indicate how a particular note is created (e.g., as a saved response, as written note which is written by a user, as a document generated from other notes, etc.).

[0148] For example, the second portion 6120 corresponds to a query section which can include one or more user interface elements for submitting or providing a query to the notebook application 3100 with respect to the source content or with respect to the plurality' of notes. For example, the second portion 6120 includes a plurality of user interface elements 6122 which correspond to suggested questions or actions that are related to the source content or plurality of notes. For example, the notebook application 3100 may be configured to generate the suggested questions or actions based on information included in the source content and / or based on the information displayed in the first portion 6110. The second portion 6120 may further include a text entry box by which a user can provide an input (e.g.. via a keyboard, via a voice input, etc.) to query the notebook application 3100, a user interface element which indicates the number of items of content which comprise the source content, etc.

[0149] In the example of FIG. 6A, the notebook application 3100 may be configured to dynamically change user interface elements in the second portion 6120 based on (in response to) the selection of one or more notes 6112 from among the plurality of notes provided in the first portion 6110. For example, as indicated in FIG. 6A one or more notes may be selected via a user input (e.g.. via a drag input, via selecting checkboxes, etc.) and in response to the selection of the one or more notes, the actions may include actions for creating content (e.g.. creating a study guide, creating an outline, creating a spreadsheet, creating a presentation, etc.), suggesting related ideas, etc., based on the selected notes 6114. For example, the notebook application 3100 may be configured to generate a note which corresponds to an outline of the content from the selected notes 6114, in response to receiving a selection of user interface element 6122a which corresponds to the action of creating an outline and saving the outline to a note. For example, the notebook application 3100 may be configured to implement one or more machine-learned models to generate the note which corresponds to the selected notes 6114, in response to receiving a selection of a user interface element which corresponds to an action of creating content with respect to the selected one or more notes and saving the content as a note.

[0150] For example, FIG. 6B illustrates a portion of a second user interface screen of a notebook application, according to one or more example embodiments of the disclosure. For example, in FIG. 6B the first portion 6210 may correspond to the first portion 6110 of FIG. 6A.

[0151] As described with respect to FIG. 6A, the notebook application 3100 may be configured to generate a note based on one or more selected notes, by implementing one or more machine-learned models, where the selected notes may correspond to source content (e.g., source content selected by a user and used as an input for generating the note). For example, the generated note may summarize or outline the notes which have been selected as described with respect to FIG. 6 A. The notebook application 3100 may be configured to generate the generated note 6214 based on the selected notes 6114. by implementing one or more machine-learned models, where the selected notes 6114 may correspond to source content (e.g., source content selected by a user and used as an input for generating the note), and in response to receiving a selection of a user interface element (e.g., user interface element 6122a) which corresponds to an action of summarizing the selected notes 6114 to the generated note 6214. FIG. 6B illustrates the generated note 6214 which has been saved to the first portion 6210 which includes one or more other notes 6212. Further, in some implementations after the generated note 6214 is saved to the first portion 6210, the plurality of user interface elements 6122 from FIG. 6 A may be configured to dynamically change back to a previous state.

[0152] In some implementations, the notebook application 3100 may be configured to enable a generated note 6214 to be exported to other applications via selection of a user interface element to send the document to another application (e.g., a word processing application, a presentation application, a spreadsheet application, a social media application, etc.). In some implementations, the notebook application 3100 may be configured to enable a generated note 6214 and / or items of content (e.g., source content 3400) to be shared with other users via selection of a user interface element to share the document and / or source content with another user.

[0153] According to examples of the disclosure, the notebook application 3100 may be configured to generate an output (e.g., an outline, a report, a summary, etc.) via one or more machine-learned models, based on source content provided to the notebook application (e.g., by the user). The notebook application 3100 may be configured to allow a user to create various projects to complete various tasks. Each project may be configured to act in a manner similar to a folder by which a user can store various information to each project. In some implementations, an individual scratchpad may correspond to or be dedicated to a particular project. In some implementations, the notebook application 3100 may be configured to receive the source content as specified by the user. The notebook application 3100 may be configured to add, delete, or modify projects according to an input receivedfrom a user. Each project may be provided a default name, a name provided by the user, or a name generated by the notebook application 3100 (e.g., via one or more machine-learned models) based on the information stored in the project (e.g., based on the source content).

[0154] Examples of the disclosure are directed to further user-facing aspects by which a user can manage content, organize content, create content, etc., via a notebook application which is configured to implement one or more machine-learned models with respect to source content selected by the user. For example, FIG. 7 illustrates examples of notebooks or projects which can be represented in a particular manner so that a user can readily understand the contents contained within the notebook or project.

[0155] In some implementations, in response to source content being provided to the notebook application 3100, the notebook application 3100 may be configured to automatically generate (e.g., using one or more machine-learned models, one or more generative machine-learned models, semantic retrieval technologies, etc.), a graphical image (e.g., an emoji, an icon, etc.) or graphical animation which corresponds to or represents the source content. In some implementations, the graphical image or graphical animation may be overlaid on a folder which is provided as a user interface element that, when selected, causes the folder to open and display the contents of the folder to the user. In addition, or alternatively, in some implementations, in response to the source content being provided to the notebook application 3100, the notebook application 3100 may be configured to automatically generate (e.g., using one or more machine-learned models, one or more generative machine- learned models, semantic retrieval technologies, etc ), a textual description (name) which corresponds to or represents the source content. The textual description may be overlaid on the folder which is provided as a user interface element that, when selected, causes the folder to open and display the contents of the folder to the user.

[0156] Referring to FIG. 7, the notebook application 3100 may have a user-specific section 7100 which stores various projects in particular folders. For example, a first folder 7110 (e.g., default folder) may be represented by a default image 7112 and have a generic name 7114 (e.g., “Default Notebook”). For example, a second folder 7120 may be represented by a graphical image 7122 and have a textual description 7124 (e.g.. “Earnings”) which is machine-learned generated and represents or corresponds to content included in the second folder 7120. For example, in response to source content being provided to the notebook application 3100, the notebook application 3100 may be configured to automatically generate (e.g., using one or more machine-learned models, one or more generative machine-learned models, semantic retrieval technologies, etc.), the graphical image 7122 which maycorrespond to an emoji, an icon, etc., which corresponds to or represents the source content. In some implementations, the graphical image 7122 may be overlaid on the second folder 7120 which is provided as a user interface element that, when selected, causes the second folder 7120 to open and display the contents of the second folder 7120 to the user. In addition, or alternatively, in some implementations, in response to the source content being provided to the notebook application 3100, the notebook application 3100 may be configured to automatically generate (e.g., using one or more machine-learned models, one or more generative machine-learned models, semantic retrieval technologies, etc ), the textual description 7124 (name) which corresponds to or represents the source content. The textual description 7124 may be overlaid on the second folder 7120 which is provided as a user interface element that, when selected, causes the second folder 7120 to open and display the contents of the second folder 7120 to the user.

[0157] Examples of the disclosure are directed to computer implemented methods for generating a document template and for providing a user interface for organizing, managing, and creating content by implementing one or more machine-learned models with respect to source content selected by a user and the generated document template. FIG. 8A illustrates a flow diagram of an example, non-limiting computer-implemented method, according to one or more example embodiments of the disclosure. FIG. 8B illustrates another flow diagram of an example, non-limiting computer-implemented method, according to one or more example embodiments of the disclosure. FIG. 9 illustrates a block diagram of a document extractor application, according to one or more example embodiments of the disclosure.

[0158] The flow diagram of FIG. 8A illustrates a method 8000 for generating a document template that can be used for organizing, managing, and creating content by implementing one or more machine-learned models with respect to training content (e.g., sample documents, training documents, etc.) selected by a user. Although shown in a particular sequence or order, unless otherwise specified, the order of the processes can be modified. Thus, the illustrated embodiments should be understood only as examples, and the illustrated processes can be performed in a different order, and some processes can be performed in parallel. Additionally, one or more processes can be omitted in various embodiments. Thus, not all processes are required in every embodiment. Other process flows are possible.

[0159] Referring to FIG. 8A, at operation 8100 the method 8000 includes a computing device receiving an input from a user relating to the selection of sample documents (e.g., training content, training documents, etc.). As described herein, the computing device may be embodied as computing device 100, server computing system 300, or combinations thereof.For example, the input may be provided by the user via input device 150. For example, the input may be provided by selecting particular files or documents which are uploaded to the computing device for use by the document extractor application 136. In some implementations, the content can be uploaded from a local memory. from another application (e.g., a portable document file application), from copied text, or from a website. The selected files, text, documents, etc. may be referred to as training content, training documents, sample documents, etc. In some implementations, the training content may be a subset of a larger corpus of content. The input may be provided or input to document extractor application 136 or document extractor application 336, for example.

[0160] In some implementations, a response to the input selecting the source content may be processed at computing device 100 without involving the server computing system 300. In some implementations, the input selecting the training content may be transmitted from computing device 100 to server computing system 300 and at least part of the response to the input may be processed by the server computing system 300. For example, the input relating to the selection of the training content may be provided at the computing device 100 and the server computing system 300 may be configured to perform an operation in response to receiving an indication of the input.

[0161] At operation 8200 the method 8000 includes the computing device receiving an input from a user requesting that a document template be generated in relation to the selection of sample documents (e.g., training content, training documents, etc.). For example, the input may be provided by the user via input device 150. In some implementations, the computing device (e.g., document extractor application 136) may be configured to provide, for presentation on a display device, a graphical user interface by which a user can request the document template to be generated. For example, the input may be provided by selecting a user interface element that is associated with generating the document template. The input may be provided or input to document extractor application 136 or document extractor application 336, for example.

[0162] At operation 8300 the method 8000 includes the computing device implementing one or more machine-learned models with respect to the selected training content (sample documents, training documents, etc.) to generate a document template. In some implementations, the document template generated by the one or more machine-learned models may include a plurality of sections, a plurality of headings that indicate different sections of the document, etc. In some implementations, the document template may furtherbe associated with a particular style, an intent, and / or a format that can be inferred or scraped from the content of the training content.

[0163] For example, the computing device can obtain information indicating that the user has selected the training content. The computing device can process the training content with one or more machine-learned models (e.g., one or more large language models) to obtain a language output. The computing device can then use the one or more machine-learned models (e.g.. one or more large language models, one or more generative machine-learned models, etc.) to generate a summarization output. In particular, a machine-learned large language model can be trained to process a variety of outputs to generate a language output. For example, the machine-learned large language model can process an embedding generated by a machine-learned embedding generation model, portions of the training content identified using the embedding generation model, language outputs generated using the machine- learned large language model or some other model, etc.

[0164] In some implementations, the one or more machine-learned models may be configured to determine (leam) an intent, style, and / or format of the training content, for example, via various natural language processing operations. For example, the training content may be broken down into tokens (e.g., words, phrases, individual characters, etc.), and converted into an embedding (e.g., numerical vector representation) which can capture semantic information regarding the training content. In some implementations, each training document among the plurality of training documents may be classified as a particular type of document (e.g., a resume, PRD, outline, legal opinion, etc ). The training document can be classified based on an aggregation of token embeddings to create a representation for the entire training document (e.g., via an averaging of the embeddings, TF-IDF weighting, etc.). The one or more machine-learned models may be configured to analyze the vocabulary in the training documents (e.g., based on the frequency of certain words, the presence of specific terms, use of domain-specific jargon, etc.). The one or more machine-learned models may also be configured to determine a syntax of a training document (e.g., based on sentence structure, sentence length, use of grammatical constructs, etc.) which can provide information regarding a particular style and / or intent of the document. The one or more machine-learned models may be configured to analyze the training content to determine semantic information based on the meaning of the content (e.g., the meaning of particular sentences and paragraphs of a document), to identify a particular style and / or intent of the training document, etc. Further, the one or more machine-learned models may be configured to determine a context of each word in relation to the entire training document.

[0165] To determine (leam) a format of the training content, the one or more machine-learned models may be configured to analyze a sequential structure of the training content to identify recurring patterns (e.g., to recognize headers, subheadings, paragraphs, bullet points, numbered lists, and other common formatting elements). The one or more machine-learned models may also be configured to leam a document structure based on consistent patterns or layouts (e.g., tables, images, captions, etc.) to understand the spatial relationships between different elements. The one or more machine-learned models may also be configured to identify specific formatting conventions (e.g., the use of indentation, font styles, font sizes, etc.), to identify the document structure and perform pattern matching. The one or more machine-learned models may be configured to leam and recognize specific document formats such that when the user uploads a plurality of training documents a document template can be generated that conforms with the document structure of the training documents. In some implementations, the training content may share common features (e.g., a common format, a common style, a common intent, etc.).

[0166] In some implementations, the one or more machine-learned models may be configured to receive information from a user which identifies information about the training content(e g., labeled data, such as an identification of the document type, the document style, the document intent, the document format, etc.). The one or more machine-learned models may be trained and / or refined based on the labeled data as well as by feedback provided via a user. The document template generated by the one or more machine-learned models (e.g., one or more large language models, one or more generative machine-learned models, etc.) at operation 8300 may be stored in the computing device and may be output for presentation on the display device 160 to the user.

[0167] The flow diagram of FIG. 8B illustrates a method 8400 for generating an output document based on the document template generated according to the flow diagram of FIG. 8A, by implementing one or more machine-learned models with respect to source content (e.g., source documents) selected by a user. Although shown in a particular sequence or order, unless otherwise specified, the order of the processes can be modified. Thus, the illustrated embodiments should be understood only as examples, and the illustrated processes can be performed in a different order, and some processes can be performed in parallel. Additionally, one or more processes can be omitted in various embodiments. Thus, not all processes are required in every' embodiment. Other process flows are possible.

[0168] Referring to FIG. 8B, at operation 8500 the method 8400 includes a computing device receiving an input from a user relating to the selection of source documents (e.g., sourcedocuments). As described herein, the computing device may be embodied as computing device 100, server computing system 300. or combinations thereof. For example, the input may be provided by the user via input device 150. For example, the input may be provided by selecting particular files or documents which are uploaded to the computing device for use by the document extractor application 136. In some implementations, the content can be uploaded from a local memory, from another application (e.g., a portable document file application), from copied text, or from a website. The selected files, text, documents, etc. may be referred to as source content, source documents, etc. In some implementations, the source content may be a subset of a larger corpus of content. The input may be provided or input to document extractor application 136 or document extractor application 336, for example.

[0169] In some implementations, a response to the input selecting the source content may be processed at computing device 100 without involving the server computing system 300. In some implementations, the input selecting the source content may be transmitted from computing device 100 to server computing system 300 and at least part of the response to the input may be processed by the server computing system 300. For example, the input relating to the selection of the source content may be provided at the computing device 100 and the server computing system 300 may be configured to perform an operation in response to receiving an indication of the input (e.g., the generation of the output document).

[0170] At operation 8600 the method 8400 includes the computing device receiving an input from a user requesting that an output document be generated in relation to the selection of the source documents (e.g., source content) and a particular document template that can also be selected via a user input. For example, the inputs may be provided by the user via input device 150. In some implementations, the computing device (e.g., document extractor application 136) may be configured to provide, for presentation on a display device, a graphical user interface by which a user can request the output document be generated in association with a particular document template that can also be selected via a user input. For example, the inputs may be provided by selecting user interface elements that are associated with selecting a desired document template and generating the output document. In some implementations, an input may be provided to generate the output document and the document extractor application 136 or document extractor application 336 may be configured to determine an applicable document template that can be applied to the source documents for generating the output document. The inputs may be provided or input to document extractor application 136 or document extractor application 336, for example.

[0171] At operation 8700 the method 8400 includes the computing device implementing one or more machine-learned models (e.g., one or more large language models, one or more generative machine-learned models, etc.) with respect to the selected source documents (source content) and an identified document template, to generate the output document. In some implementations, the output document generated by the one or more machine-learned models may include a plurality of sections, a plurality of headings that indicate different sections of the document, etc., which are in conformance with the identified or selected document template. In some implementations, the output document may further be associated with a particular sty le, an intent, and / or a format that is based on the sty le, intent, and / or format of the document template.

[0172] For example, the computing device can obtain information indicating that the user has selected the source documents. The computing device can process the source documents with one or more machine-learned models (e.g., one or more large language models) to obtain a language output. The computing device can then use the one or more machine- learned models (e.g.. one or more large language models, one or more generative machine- learned models, etc.) to generate a summarization output. In particular, a machine-learned large language model can be trained to process a variety' of outputs to generate a language output. For example, the machine-learned large language model can process an embedding generated by a machine-learned embedding generation model, portions of the source documents identified using the embedding generation model, language outputs generated using the machine-learned large language model or some other model, etc.

[0173] In some implementations, the one or more machine-learned models may be configured to determine (learn) an intent, sty le, and / or format of the source documents, for example, via various natural language processing operations. For example, the source documents may be broken down into tokens (e.g., words, phrases, individual characters, etc.), and converted into an embedding (e.g., numerical vector representation) which can capture semantic information regarding the source documents. In some implementations, each source document among the plurality of source documents may be classified as a particular type of document (e.g., a resume, PRE), outline, legal opinion, etc.). The source document can be classified based on an aggregation of token embeddings to create a representation for the entire training document (e.g., via an averaging of the embeddings, TF-IDF weighting, etc.). The one or more machine-learned models may be configured to analyze the vocabulary in the source documents (e.g., based on the frequency of certain words, the presence of specific terms, use of domain-specific jargon, etc.). The one or more machine-learned models may also beconfigured to determine a syntax of a source document (e.g., based on sentence structure, sentence length, use of grammatical constructs, etc.) which can provide information regarding a particular style and / or intent of the source document. The one or more machine-learned models may be configured to analyze the source documents to determine semantic information based on the meaning of the content (e.g., the meaning of particular sentences and paragraphs of a document), to identify a particular style and / or intent of the source document, etc. Further, the one or more machine-learned models may be configured to determine a context of each word in relation to the entire source document.

[0174] The one or more machine-learned models (e.g., one or more large language models, one or more generative machine-learned models, etc.) may be configured to apply the document template to the plurality of source documents to generate the output document. For example, the one or more machine-learned models may be configured to extract first content (e.g., background information, title information, body information, conclusion information, etc.) from the plurality of source documents and associate the first content with a first section of the output document (e.g., a background section, a title section, a body section, a conclusion section, etc.). Likewise, the one or more machine-learned models may be configured to extract second content from the plurality of source documents and associate the second content with a second section of the output document, and so on. The one or more machine-learned models (e.g., one or more large language models, one or more generative machine-learned models, etc.) may be configured to associate content from the source documents with a particular section of the output document based on the determined semantic information, context information, intent information, etc., that is associated with that content. For example, if the one or more machine-learned models determines (e.g., based on a confidence level) that a certain portion of a source document is associated with background information regarding a certain topic or subject, the one or more machine-learned models may be configured to implement some or all of the certain portion in a background section of the output document.

[0175] For example, the one or more machine-learned models may be configured to apply the style, intent, and / or format of the document template to the content which is extracted from the plurality of source documents. For example, if the document template is associated with a persuasive intent and opinionated style, the one or more machine-learned models may be configured to generate the output document with such features based on the content from the plurality of source documents, where the output document may have a document structure that is defined by or associated with the document template.

[0176] As another example, the one or more machine-learned models may be configured to identify a document type associated with the plurality of source documents based on the content of each of the plurality of source documents. The one or more machine-learned models may be configured to apply the style, intent, and / or format of a document template which is associated with the identified document type. For example, if the one or more machine-learned models determines the document type associated with the source documents is a PRD, the one or more machine-learned models may be configured to apply the style, intent, and / or format of a document template which is associated with the PRD document type to the content from the plurality of source documents.

[0177] In some implementations, the one or more machine-learned models may be configured to receive information from a user which identifies information about the output document and / or the source documents (e.g., labeled data, such as an identification of the document type, the document style, the document intent, the document format, etc.). The one or more machine-learned models for generating the output document may be trained and / or refined based on the labeled data as well as by feedback provided via a user. For example, the output document generated at operation 8700 may be stored in the computing device and may be output for presentation on the display device 160 to the user.

[0178] Referring to FIG. 9, the document extractor application 9100 (which may correspond to document extractor application 136 and / or document extractor application 336) may include a conditioning parameters generator 9110. one or more sequence processing models 9120, one or more large language models 9130, and one or more generative machine-learned models 9140. The document extractor application 9100 may receive an input 9200 from a user as discussed above with respect to operations 8100, 8200, 8500, 8600 of FIGS. 8A and 8B. Conditioning parameters generator 9110 may be configured to generate conditioning parameters based at least in part on the input, wherein the conditioning parameters provide values for one or more conditions associated with content to be generated which relates at least in part to the input 9200 and training content 9400 and / or source content 9500 selected by the user.

[0179] For example, the training content 9400 and source content 9500 can include any kind of document (e.g., in digital form) and may include books, product manuals, legal opinions, academic papers, proprietary data files, patent documents, web pages, emails, forum posts, social media posts, videos, images, geographic information, or any other type or manner of content which may be stored or accessed in digital form (e.g.. in a database, memory device, etc.). In some implementations, the training content 9400 and source content 9500 may bestored in the content data store 350 by the user selecting certain documents, images, or other content to store in the content data store 350. In some implementations, the training content 9400 and source content 9500 may be stored at the computing device 100 and / or server computing system 300.

[0180] To generate the conditioning parameters, the conditioning parameters generator 9110 may be configured to retrieve values for the one or more conditions associated with the input. For example, to generate the conditioning parameters, the conditioning parameters generator 9110 may be configured to extract the values for the one or more conditions from the input. The input may include information indicative of the user's intent or requirements. In some implementations, the conditioning parameters generator 9110 (or the one or more sequence processing models 9120 or the one or more large language models 9130) may be configured to extract information from the input 9200 to identify values for the one or more conditions, and the conditioning parameters generator 9110 may be configured to generate the conditioning parameters based on the extracted values. For example, the input itself may identify a color to be used for headings in a generated document (e.g., “blue font for the title”) or an attribute or feature (e.g.. “circle bullet points”) that can be used to generate the conditioning parameters for generating a document template related to the training content 9400 or for generating an output document related to the source content 9500.

[0181] To generate the conditioning parameters, the conditioning parameters generator 9110 may be configured to infer the values for the one or more conditions from the input. The input may include information indicative of the user's intent or requirements. In some implementations, the conditioning parameters generator 9110 (or the one or more sequence processing models 9120 or the one or more large language models 9130) may be configured to infer information from the input 9200 to identify values for the one or more conditions, and the conditioning parameters generator 9110 may be configured to generate the conditioning parameters based on the inferred values. For example, the input may include a reference to a length (“short,” “long,” etc.) of the output document to be generated based on the source content 9500, and the conditioning parameters generator 9110 (or the one or more sequence processing models 9120 or the one or more large language models 9130) may be configured to infer a value based on the input. For example, an input requesting the document extractor application 9100 to generate a “standard” resume may infer a value of about 1 page while a “long” resume may be associated with a value of about 2 to 3 pages. For example, the document extractor application 9100 may be configured to ascertain aninferred value based on information via external content 9300 (e.g., a website which describes lengths of resumes).

[0182] In some implementations, the conditioning parameters generator 9110 may be configured to infer the values for the one or more conditions from the input by providing the input to one or more sequence processing models 9120, wherein the one or more sequence processing models 9120 are configured to output the values for the one or more conditions in response to or based on the query. The one or more sequence processing models 9120 may include one or more machine-learned models which are configured to process and analyze sequential data and to handle data that occurs in a specific order or sequence, including time series data, natural language text, or any other data with a temporal or sequential structure.

[0183] The one or more sequence processing models 9120 may receive an input including text and tokenize the input by breaking down the sequence of text into small units (tokens) to provide a structured representation of the input sequence. The one or more sequence processing models 9120 may represent the tokens as vectors in a continuous vector space bymapping each token to a high-dimensional vector, where the relationships between tokens (words) are reflected in the geometric relationships between their corresponding vector. For example, the one or more sequence processing models 9120 may receive an input extracted from the training content 9400 and / or source content 9500 including the text “the bustling marketplace'’ and tokenize the input by breaking down the sequence of text into small units (tokens) (e.g., “the,” “bustling,” and “marketplace”), thereby providing a structured representation of the input sequence. In a word embedding, semantically similar words are closer together in the vector space. For example, the vectors for "bustling" and "busy" might be close to each other because of their semantic relationship, while the vectors for “bustling” and “stagnant” may be far apart compared to the vectors for “bustling” and “stagnant”.

[0184] The one or more large language models 9130 can be, or otherwise include, a model that has been trained on a large corpus of language training data in a manner that provides the one or more large language models 9130 with the capability to perform multiple language tasks. For example, the one or more large language models 9130 can be trained to perform summarization tasks, conversational tasks, simplification tasks, oppositional viewpoint tasks, etc. In particular, the one or more large language models 9130 can be trained to process a variety7of outputs to generate a language output. For example, the one or more large language models 9130 can process an embedding generated by a machine-learned embedding generation model, portions of source content or training content (e.g., document chunk(s))identified using an embedding generation model, language outputs generated using the one or more large language models 9130 or some other model, etc.

[0185] The one or more generative machine-learned models 9140 may include a deep neural network or a generative adversarial network (GAN), variational autoencoders, stable diffusion machine-learned models, visual transformers, neural radiance fields (NeRFs), etc., to generate content (e.g.. a resume, an outline, a PRD. etc.) with values for conditions associated with one or more features. For example, the computing device may include a database (e.g., machine-learned model data store 370) which is configured to store a plurality of generative machine-learned models respectively associated with a plurality' of different types of content or a plurality' of different types of documents (e.g., different genres or subjects, different kinds of content including imagery, videos, and text, different types of content including outlines, reports, spreadsheets, resumes, PRDs, etc.).

[0186] In some implementations, the computing device may be configured to retrieve, from among the one or more generative machine-learned models 9140, a generative machine- learned model associated with a particular type of content (document) and / or document template, for generating the output document, relating to the input. In some implementations, the computing device may be configured to retrieve, from among the one or more generative machine-learned models 9140, a generative machine-learned model associated with a particular type of content (document) for generating a particular type of document template, relating to the input.

[0187] In some implementations, the one or more generative machine-learned models 9140 may be trained on a large dataset of content (e.g., a large corpus of language training data) with corresponding information about the conditions associated with the content. During training, the one or more generative machine-learned models 9140 may be configured to learn relationships between elements in an output (e.g., content) and conditions that influence them. This may involve the computing device adjusting each generative machine-learned model’s internal parameters to generate realistic or accurate content (e.g., grammatically correct content, coherent content, etc.) based on the training data. The one or more generative machine-learned models 9140 may be trained on one or more training datasets including a plurality of reference document templates. The one or more generative machine- learned models 9140 may be trained on one or more training datasets including a plurality of reference output documents that are associated with one or more document templates. The one or more training datasets may include values for the one or more conditions.

[0188] In some implementations, the one or more generative machine-learned models 9140 are configured to generate the document template 9700 in response to receiving the selection of the training content 9400. For example, the document template 9700 may be generated based on the conditioning parameters (and corresponding values for the one or more conditions) to make decisions for generating the content of the document template 9700. In some implementations, the one or more generative machine-learned models 9140 are configured to generate the output document 9800 in response to receiving the selection of the source content 9500 and based on a particular document template 9700 which may be selected by the user or may be automatically determined based on the source content 9500. For example, the output document 9800 may be generated based on the conditioning parameters (and corresponding values for the one or more conditions) to make decisions for generating the content of the output document 9800.

[0189] In some implementations, the server computing system 300 may provide (transmit) content or a portion of the generated content to computing device 100 or the server computing system 300 may provide access to the generated content to the computing device 100. For example, the document template 9700 and / or the output document 9800 may be generated at the server computing system 300 and stored at one or more computing devices (e.g., one or more of computing device 100, external computing device 200, server computing system 300, external content 500, content data store 350, user data store 360, etc.).

[0190] In some implementations, after the document template 9700 is generated, the user can provide feedback or a further input relating to the document template 9700 which is generated based on the training content 9400 provided (and / or based on a query provided via the user), and one or more of the operations 8100 through 8300 can be repeated. In some implementations, after the output document 9800 is generated, the user can provide feedback or a further input relating to the output document 9800 which is generated based on the source content 9500 provided (and / or based on a query provided via the user), and one or more of the operations 8500 through 8700 can be repeated.

[0191] Examples of the disclosure are also directed to user-facing aspects by which a user can manage content, organize content, create content, etc., via a notebook application and / or document extractor application which are each configured to implement one or more machine-learned models with respect to content selected by the user. For example, FIGS. 10A through 10F illustrate examples of actions which can be implemented for a project in which a document template is generated via one or more machine-learned models based on training content selected by a user, and in which an output document is generated via one ormore machine-learned models based on source content selected by the user, according to one or more example embodiments of the disclosure.

[0192] For example, FIG. 10A illustrates a first user interface screen (e.g., atemplate builder user interface screen) of a document extractor application (which may be incorporated as part of a notebook application), according to one or more example embodiments of the disclosure.

[0193] In FIG. 10A. the first user interface screen 1010 depicts a user interface (e.g., a template builder user interface screen) which provides information about the document extractor application 9100. In particular, document extractor application 9100 is configured to present for display the first user interface screen 1010 which includes a visual depiction 1012 regarding how a document template can be created and structured based on training content and information 1014 regarding features about the document extractor application 9100.

[0194] As illustrated in FIG. 10A, the first user interface screen 1010 further includes a plurality7of user interface elements which are selectable (or can be interacted with) by the user for generating a document template. For example, a first portion of the first user interface screen 1010 includes a plurality of first user interface elements 1016 are configured to be selectable as training content for creating a document template. In FIG. 10A, the training content which can be selected by the user for generating a document template includes a PRD file (“Product PRD”), notes from a meeting (“Brainstorm meeting"’), and a note regarding strategy (“Notebook LLM Strategy7”).

[0195] For example, a second portion of the first user interface screen 1010 includes a second user interface element 1018 which is configured to enable a user to upload one or more training documents for generating a document template. As illustrated in FIG. 10 A, a third user interface element 1019 is configured to receive an input associated with creating (generating) a document template, based on the training content (e.g., training documents) identified (selected) by the user via the first user interface screen 1010. For example, the document extractor application 9100 may be configured to generate the document template according to the examples described herein based on the training content selected by the user and in response to the user input (e.g.. selecting the third user interface element 1019).

[0196] FIG. 10B illustrates a second user interface screen (e.g., a document template customization user interface screen) of a document extractor application (which may be incorporated as part of a notebook application), according to one or more example embodiments of the disclosure.

[0197] In FIG. 10B, the second user interface screen 1020 depicts a user interface (e.g., a template builder user interface screen) which is associated with enabling a user to customize one or more features associated with the document template generated by the document extractor application 9100. In particular, document extractor application 9100 is configured to present for display the second user interface screen 1020 which includes various portions and user interface elements by which the user can modify or customize a document template generated by the document extractor application 9100.

[0198] For example, the second user interface screen 1020 includes a first portion 1021 which identifies the name of the generated document template C'PRD template’’).

[0199] For example, the second user interface screen 1020 includes a second portion 1022 which is associated with the training documents that were selected for generating the document template. For example, training documents 1023 include the “Product PRD” training document and the “Brainstorm meeting” document. A first user interface element1024 may be configured to enable a user to add training documents so that the document template can be re-generated based on the added training documents.

[0200] For example, the second user interface screen 1020 includes a third portion 1025 which is associated with a st le of the generated document template. For example, the third portion1025 of the second user interface screen 1020 includes a plurality7of first user interface elements 1026 which are associated with different possible styles that can be selected for generating (or re-generating) the document template. In FIG. 10B, example styles which can be selected include an “Expert style voice”, “Casual language”, “Uses metaphors”, “Opinionated”, and “MBT Type: INTJ”. A second user interface element 1027 may be configured to enable a user to add a new style so that the document template can be regenerated based on the added sty le(s).

[0201] For example, the second user interface screen 1020 includes a fourth portion 1028 which is associated with a document format of the generated document template. For example, the fourth portion 1028 of the second user interface screen 1020 includes a plurality7of second user interface elements 1029 which are associated with different sections of a document structure that can be selected and modified for generating (or re-generating) the document template. In FIG. 10B, example document sections which are visible in the drawing and which can be selected include a “Title” section and a “Description” section. Other document sections may include a “Background” section and a “Conclusion” section, for example. Each of the plurality of second user interface elements 1029 may be configured to be manipulated by a user such that the different sections can be rearranged according to auser input, so that the document template can be re-generated based on the rearranged document structure (format).

[0202] Though not shown in FIG. 10B, the second user interface screen 1020 can also include a further portion which is associated with an intent of the generated document template. For example, the further portion of the second user interface screen 1020 can include a plurality7of user interface elements which are associated with different possible intents that can be selected for generating (or re-generating) the document template. Example intent which can be selected include a “Persuasive” intent, an “Informative” intent, an “Entertain” intent, an “Inspire” intent, and the like.

[0203] FIG. 10C illustrates a visual depiction of how the document extractor application (which may be incorporated as part of a notebook application) can leam a format, style, and / or intent of a training document, according to one or more example embodiments of the disclosure. As illustrated in FIG. 10C, a training document 1032 (“PRD for Product”) includes a plurality of sections 1034 for a PRD document, including an introduction section, critical user journey (CUJ) section, target audience section, problem statement section, and a proposal section. The document extractor application 9100 may be configured to leam the document type of the training document 10323 via selection of a user interface element 1036. As shown in FIG. 10C, the learned document sections 1038 for a PRD document include the same sections as the training document 1032. The document extractor application 9100 may be configured to generate a PRD in response to a user uploading source documents, where the generated PRD may have the format shown in FIG. 10C or a similar format, based on the learned document type information which is obtained from the training document 1032.

[0204] For example. FIG. 10D illustrates a fourth user interface screen (e.g., an output document generation user interface screen) of a document extractor application (which may be incorporated as part of a notebook application), according to one or more example embodiments of the disclosure.

[0205] In FIG. 10D, the fourth user interface screen 1040 depicts a user interface (e.g., an output document generation user interface screen) which enables a user to create (generate) an output document based on one or more source documents which can be selected by a user, according to a document template that can also be selected or provided to the user, for example, based on the content of the source documents. For example, in FIG. 10D, the user has selected a plurality of source documents 1042 (which may correspond to notes that are added to the scratchpad or notes section in the notebook application 3100).

[0206] In some implementations, the user may identify or select a particular document template and provide an input that causes the document extractor application 9100 to generate an output document based on the selected source documents and the selected document template. In some implementations, the document extractor application 9100 may be configured to analyze the content of the selected source documents and determine or suggest one or more document templates which may be appropriate or applicable to the source documents.

[0207] For example, in FIG. 10D, the fourth user interface screen 1040 includes a first user interface element 1044 which is configured to, when selected, create or generate a PRD based on the document template that is previously generated and associated with PRD documents.

[0208] For example, FIG. 10E illustrates a fifth user interface screen (e.g.. an output document user interface screen) of a document extractor application (which may be incorporated as part of a notebook application), according to one or more example embodiments of the disclosure.

[0209] In FIG. 10E, the fifth user interface screen 1050 depicts a user interface (e.g., an output document user interface screen) which includes the output document generated by the document extractor application 9100 according to the selected source documents and the document template (e g., the PRD document template). For example, the output document may correspond to a note 1052 for the notebook application 3100. As illustrated in FIG. 10E, the output document may have a document structure or format that is consistent with a format of a PRD document, and may include similar sections such as an introduction section 1054 and a CUJs section 1056. The document extractor application 9100 may be configured to generate content for each section based on the content of the source documents, via one or more machine-learned models, as described according to the examples provided herein. For example, the output document may be stored in the computing device, may be transmitted to another computing device, may be saved as a particular document file type in another document application (e.g., via first user interface element 1058), may be shared with another user (e.g., via second user interface element 1059), etc.

[0210] FIG. 11A depicts a block diagram of an example computing system for organizing, managing, and creating content by implementing one or more machine-learned models with respect to source content and / or training content selected by a user, according to one or more example embodiments of the disclosure. The system 1100 includes a user computing device 1102, a server computing system 1130, and a training computing system 1150 that are communicatively coupled over a network 1180.

[0211] FIG. 1 IB depicts a block diagram of an example computing device for organizing, managing, and creating content by implementing one or more machine-learned models with respect to source content and / or training content selected by a user, according to one or more example embodiments of the disclosure.

[0212] FIG. 11C depicts a block diagram of an example computing device for organizing, managing, and creating content by implementing one or more machine-learned models with respect to source content and / or training content selected by a user, according to one or more example embodiments of the disclosure.

[0213] The user computing device 1102 (which may correspond to computing device 100) can be any type of computing device, such as, for example, a personal computing device (e.g., laptop or desktop), a mobile computing device (e.g., smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0214] The user computing device 1102 includes one or more processors 1112 and a memory 1114. The one or more processors 1112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA. a controller, a microcontroller, etc.) and can be one processor or a plurality’ of processors that are operatively connected. The memory 1114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 1114 can store data 1116 and instructions 1118 which are executed by the processor 1 112 to cause the user computing device 1 102 to perform operations.

[0215] In some implementations, the user computing device 1102 can store or include one or more machine-learned models 1120 (e.g., large language models, sequence processing models, generative machine-learned models, etc.). For example, the one or more machine- learned models 1120 can be or can otherwise include various machine-learned models such as neural networks (e.g., deep neural networks) or other types of machine-learned models, including non-linear models and / or linear models. Example neural networks can include feed-forward neural networks, recurrent neural networks (RNNs), including long short-term memory (LSTM) based recurrent neural networks, convolutional neural networks (CNNs), diffusion models, generative-adversarial networks, or other forms of neural networks. Example neural networks can be deep neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine-learned models can include multi-headed self-attention models (e.g.,transformer models). Example machine-learned models were described herein with reference to FIGS. 1 A through 10E.

[0216] In some implementations, the one or more machine-learned models 1 120 can be received from the server computing system 1130 over network 1180, stored in the memon1114, and then used or otherwise implemented by the one or more processors 1112. In some implementations, the user computing device 1102 can implement multiple parallel instances of a single machine-learned model (e.g.. to perform parallel tasks across multiple instances of the machine-learned model). In some implementations, the task is a generative task and one or more machine-learned models may be implemented to output content (e.g., a document template, an output document, etc.) in view of various inputs (e.g., a query, training documents, source documents, conditioning parameters, etc.). More particularly, the machine-learned models disclosed herein (e.g., including large language models, sequence processing models, generative machine-learned models, etc.), may be implemented to perform various tasks related to an input query.

[0217] According to examples of the disclosure, a computing system may implement one or more sequence processing models 3120. 9120 as described herein to output values for the one or more conditions in response to or based on the query. The one or more sequence processing models 3120, 9120 may include one or more machine-learned models which are configured to process and analyze sequential data and to handle data that occurs in a specific order or sequence, including time series data, natural language text, or any other data with a temporal or sequential structure.

[0218] According to examples of the disclosure, a computing system may implement one or more large language models 3130, 9130 to determine a plurality of variables based on the query. For example, a large language model may include a Bidirectional Encoder Representations from Transformers (BERT) large language model. The large language model may be trained to understand and process natural language for example. The large language model may be configured to extract information from the input (e.g., a query, training documents, source documents, etc.) to identify keywords, intents, and context within the input to determine a plurality of variables for generating content. The variables may include latent variables that represent an underlying structure of the language.

[0219] According to examples of the disclosure, a computing system may implement one or more generative machine-learned models 3140, 9140 to generate various content (e.g., for generating an outline, a summary, a response to a query, a document template, an output document generated based on the document template, etc.) having values for one or moreconditions. The one or more generative machine-learned models 3140, 9140 may include a deep neural network or a generative adversarial network (GAN) to generate the content with one or more features having values for one or more conditions associated with the features. For example, the one or more generative machine-learned models 3140, 9140 may include variational autoencoders, stable diffusion machine-learned models, visual transformers, neural radiance fields (NeRFs), etc., to generate the content.

[0220] Additionally, or alternatively, one or more machine-learned models 1140 can be included in or otherwise stored and implemented by the server computing system 1130 that communicates with the user computing device 1102 according to a client-server relationship. For example, the one or more machine-learned models 1140 can be implemented by the server computing system 1130 as a portion of a web service (e.g.. a navigation service, a word processing sendee, an educational service, and the like). Thus, one or more machine- learned models 1120 can be stored and implemented at the user computing device 1102 and / or one or more machine-learned models 1140 can be stored and implemented at the sen’ er computing system 1130.

[0221] The user computing device 1102 can also include one or more user input components 1122 that receives user input. For example, the user input component 1122 can be a touch- sensitive component (e.g., a touch-sensitive display screen or a touch pad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component can serve to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other devices and methods by which a user can provide a user input.

[0222] The server computing system 1130 (which may correspond to server computing system 300) includes one or more processors 1132 and a memory 1134. The one or more processors 1 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 1134 can include one or more non-transitory computer-readable storage media, such as RAM. ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 1134 can store data 1136 and instructions 1138 which are executed by the processor 1132 to cause the server computing system 1130 to perform operations.

[0223] In some implementations, the server computing system 1130 includes or is otherwise implemented by one or more server computing devices. In instances in which the server computing system 1130 includes a plurality of server computing devices, such servercomputing devices can operate according to sequential computing architectures, parallel computing architectures, or some combination thereof.

[0224] As described above, the server computing system 1130 can store or otherwise include one or more machine-learned models 1140. For example, the one or more machine-learned models 1140 can be or can otherwise include various machine-learned models. Example machine-learned models include neural networks or other multi-layer non-linear models. Example neural networks can include feed-forward neural networks, recurrent neural networks (RNNs), including long short-term memory (LSTM) based recurrent neural networks, convolutional neural networks (CNNs), diffusion models, generative-adversarial networks, or other forms of neural networks. Example neural networks can be deep neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine-learned models can include multiheaded self-attention models (e.g., transformer models). Example machine-learned models were described herein with reference to FIGS. 1 A through 10E.

[0225] The user computing device 1102 and / or the server computing system 1130 can train the one or machine-learned models 1120 and / or 1140 via interaction with the training computing system 1150 that is communicatively coupled over the network 1180. The training computing system 1150 can be separate from the server computing system 1130 or can be a portion of the server computing system 1130.

[0226] The training computing system 1150 includes one or more processors 1152 and a memory 1 154. The one or more processors 1 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 1154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 1154 can store data 1156 and instructions 1158 which are executed by the processor 1152 to cause the training computing system 1150 to perform operations. In some implementations, the training computing system 1150 includes or is otherwise implemented by one or more server computing devices.

[0227] The training computing system 1150 can include a model trainer 1160 that trains the one or more machine-learned models 1120 and / or 1140 stored at the user computing device 1102 and / or the server computing system 1130 using various training or learning techniques, such as, for example, backwards propagation of errors. For example, a loss function can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g..based on a gradient of the loss function). Various loss functions can be used such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques can be used to iteratively update the parameters over a number of training iterations.

[0228] In some implementations, performing backwards propagation of errors can include performing truncated backpropagation through time. The model trainer 1160 can perform a number of generalization techniques (e.g.. weight decays, dropouts, etc.) to improve the generalization capability of the models being trained.

[0229] In particular, the model trainer 1160 can train the one or more machine-learned models 1120 and / or 1140 based on a set of training data 1162. The training data 1162 can include, for example, various datasets which may be stored remotely or at the training computing system 1150. For example, in some implementations an example dataset utilized for training includes a large corpus of language training data that provides one or more large language models with the capability to perform multiple language tasks. For example, the one or more large language models can be trained to perform summarization tasks, conversational tasks, simplification tasks, oppositional viewpoint tasks, etc. In particular, the one or more large language models can be trained to process a variety of outputs to generate a language output. However, other datasets (e.g., of images) may be utilized (e.g., images obtained from external websites). In some implementations, the dataset may be confined to a particular genre or subject, particular kinds of content including imagery, videos, and text, particular styles or types of content (e.g., outlines, reports, presentations, spreadsheets, resumes, PRDs, etc ), etc. In some implementations, the dataset may contain diverse subject matter.

[0230] In some implementations, if the user has provided consent, the training examples can be provided by the user computing device 1102. Thus, in such implementations, the one or more machine-learned models 1120 provided to the user computing device 1102 can be trained by the training computing system 1150 on user-specific data received from the user computing device 1102. In some instances, this process can be referred to as personalizing the model.

[0231] The model trainer 1160 includes computer logic utilized to provide desired functionality. The model trainer 1160 can be implemented in hardware, firmware, and / or software controlling a general purpose processor. For example, in some implementations, the model trainer 1160 includes program files stored on a storage device, loaded into a memory and executed by one or more processors. In other implementations, the model trainer 1160includes one or more sets of computer-executable instructions that are stored in a tangible computer-readable storage medium such as RAM, hard disk, or optical or magnetic media.

[0232] The network 1180 can be any type of communications network, such as a local area network (e.g., intranet), wide area network (e.g., Internet), or some combination thereof and can include any number of wired or wireless links. In general, communication over the network 1180 can be carried via any t pe of wired and / or wireless connection, using a wide variety of communication protocols (e.g., TCP / IP. HTTP. SMTP. FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).

[0233] The machine-learned models described in this specification may be used in a variety of tasks, applications, and / or use cases.

[0234] In some implementations, the input to the machine-learned model(s) of the disclosure can be text or natural language data. The machine-learned model(s) can process the text or natural language data to generate an output. As an example, the machine-learned model(s) can process the natural language data to generate a language encoding output. As another example, the machine-learned model(s) can process the text or natural language data to generate a latent text embedding output. As another example, the machine-learned model(s) can process the text or natural language data to generate a translation output. As another example, the machine-learned model(s) can process the text or natural language data to generate a classification output. As another example, the machine-learned model(s) can process the text or natural language data to generate a textual segmentation output. As another example, the machine-learned model(s) can process the text or natural language data to generate a semantic intent output. As another example, the machine-learned model(s) can process the text or natural language data to generate an upscaled text or natural language output (e.g., text or natural language data that is higher quality than the input text or natural language, etc.). As another example, the machine-learned model(s) can process the text or natural language data to generate a prediction output.

[0235] In some implementations, the input to the machine-learned model(s) of the disclosure can be speech data. The machine-learned model(s) can process the speech data to generate an output. As an example, the machine-learned model(s) can process the speech data to generate a speech recognition output. As another example, the machine-learned model(s) can process the speech data to generate a speech translation output. As another example, the machine- learned model(s) can process the speech data to generate a latent embedding output. As another example, the machine-learned model(s) can process the speech data to generate an encoded speech output (e.g., an encoded and / or compressed representation of the speech data,etc.). As another example, the machine-learned model(s) can process the speech data to generate an upscaled speech output (e.g., speech data that is higher quality than the input speech data, etc.). As another example, the machine-learned model(s) can process the speech data to generate a textual representation output (e.g., a textual representation of the input speech data, etc.). As another example, the machine-learned model(s) can process the speech data to generate a prediction output.

[0236] In some implementations, the input to the machine-learned model(s) of the disclosure can be sensor data. The machine-learned model(s) can process the sensor data to generate an output. As an example, the machine-learned model(s) can process the sensor data to generate a recognition output. As another example, the machine-learned model(s) can process the sensor data to generate a prediction output. As another example, the machine-learned model(s) can process the sensor data to generate a classification output. As another example, the machine-learned model(s) can process the sensor data to generate a segmentation output. As another example, the machine-learned model(s) can process the sensor data to generate a visualization output. As another example, the machine-learned model(s) can process the sensor data to generate a diagnostic output. As another example, the machine-learned model(s) can process the sensor data to generate a detection output.

[0237] FIG. 11 A illustrates an example computing system that can be used to implement aspects of the disclosure. Other computing systems can be used as well. For example, in some implementations, the user computing device 1102 can include the model trainer 1160 and the training data 1 1 2. In such implementations, the one or more machine-learned models 1120 can be both trained and used locally at the user computing device 1102. In some of such implementations, the user computing device 1102 can implement the model trainer 1160 to personalize the one or more machine-learned models 1120 based on userspecific data.

[0238] FIG. 11B depicts a block diagram of an example computing device for organizing, managing, and creating content by implementing one or more machine-learned models with respect to training content and / or source content selected by a user, according to one or more example embodiments of the disclosure. The computing device 1200 can be a user computing device or a server computing device.

[0239] The computing device 1200 includes a number of applications (e.g., applications 1 through N). Each application contains its own machine learning library and machine-learned model(s). For example, each application can include a machine-learned model. Example applications include the notebook application as described herein, the document extractorapplication as described herein, a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, a social media application, a map application, a navigation application, etc.

[0240] As illustrated in FIG. 11 B, each application can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g.. a public API). In some implementations, the API used by each application is specific to that application.

[0241] FIG. 11C depicts a block diagram of an example computing device for organizing, managing, and creating content by implementing one or more machine-learned models with respect to training content and / or source content selected by a user, according to one or more example embodiments of the disclosure. The computing device 1300 can be a user computing device or a server computing device.

[0242] The computing device 1300 includes a number of applications (e.g., applications 1 through N). Each application is in communication with a central intelligence layer. Example applications include the notebook application as described herein, the document extractor application as described herein, a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, a map application, a social media application, a navigation application, a social media application, etc. In some implementations, each application can communicate with the central intelligence layer (and model(s) stored therein) using an API (e.g., a common API across all applications).

[0243] The central intelligence layer includes a number of machine-learned models. For example, as illustrated in FIG. 11C, a respective machine-learned model can be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine-learned model. For example, in some implementations, the central intelligence layer can provide a single model for all of the applications. In some implementations, the central intelligence layer is included within or otherwise implemented by an operating system of the computing device 1300.

[0244] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for the computing device 1300. As illustrated in FIG. 11C, the central device data layer can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. Insome implementations, the central device data layer can communicate with each device component using an API (e.g.. a private API).

[0245] To the extent alleged generic terms including "module", and "unit," and the like are used herein, these terms may refer to, but are not limited to, a software or hardware component or device, such as a Field Programmable Gate Array (FPGA) or Application Specific Integrated Circuit (ASIC), which performs certain tasks. A module or unit may be configured to reside on an addressable storage medium and configured to execute on one or more processors. Thus, a module or unit may include, by way of example, components, such as software components, object-oriented software components, class components and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables. The functionality provided for in the components and modules / units may be combined into fewer components and modules / units or further separated into additional components and modules.

[0246] Aspects of the above-described example embodiments may be recorded in non- transitory computer-readable media including program instructions to implement various operations embodied by a computer. The media may also include, alone or in combination with the program instructions, data files, data structures, and the like. Examples of non- transitory computer-readable media include magnetic media such as hard disks, floppy disks, and magnetic tape; optical media such as CD ROM disks. Blue-Ray disks, and DVDs; magneto-optical media such as optical discs; and other hardware devices that are specially configured to store and perform program instructions, such as semiconductor memory, readonly memory (ROM), random access memory (RAM), flash memory, USB memory, and the like. Examples of program instructions include both machine code, such as produced by a compiler, and files containing higher level code that may be executed by the computer using an interpreter. The program instructions may be executed by one or more processors. The described hardware devices may be configured to act as one or more software modules in order to perform the operations of the above-described embodiments, or vice versa. In addition, a non-transitory computer-readable storage medium may be distributed among computer systems connected through a network and computer-readable codes or program instructions may be stored and executed in a decentralized manner. In addition, the non- transitory computer-readable storage media may also be embodied in at least one application specific integrated circuit (ASIC) or Field Programmable Gate Array (FPGA).

[0247] Each block of the flowchart illustrations may represent a unit, module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of order. For example, two blocks shown in succession may in fact be executed substantially concurrently (simultaneously) or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved.

[0248] While the disclosure has been described with respect to various example embodiments, each example is provided by way of explanation, not limitation of the disclosure. Those skilled in the art, upon attaining an understanding of the foregoing, can readily produce alterations to, variations of, and equivalents to such embodiments. Accordingly, the disclosure does not preclude inclusion of such modifications, variations and / or additions to the disclosed subject matter as would be readily apparent to one of ordinary skill in the art. For example, features illustrated or described as part of one embodiment can be used with another embodiment to yield a still further embodiment. Thus, it is intended that the disclosure covers such alterations, variations, and equivalents.

Claims

WHAT IS CLAIMED IS:

1. A computing device for generating an output document, comprising: one or more memories configured to store instructions; and one or more processors configured to execute the instructions to perform operations, the operations comprising: receiving a pl urality of source documents, receiving an input associated with a request to generate the output document based on the plurality of source documents and a particular document template, and generating, via one or more machine-learned models, the output document having the particular document template.

2. The computing device of claim 1, wherein the operations further comprise: providing, for presentation on a display device, a graphical user interface configured to receive a selection of the plurality of source documents.

3. The computing device of claim 1, wherein the operations further comprise: determining, by the one or more machine-learned models, an intent associated with the output document, based on a content associated with each of the plurality of source documents.

4. The computing device of claim 1, wherein the operations further comprise: determining, by the one or more machine-learned models, a style associated with the output document, based on a content of each of the plurality of source documents.

5. The computing device of claim 1, wherein the operations further comprise: determining, by the one or more machine-learned models, a format associated with the output document, based on a content of each of the plurality of source documents.

6. The computing device of claim 5, wherein determining, by the one or more machine-learned models, the format associated with the output document, based on the content of each of the plurality of source documents, includes identifying a plurality of headings for respective sections of the output document.

7. The computing device of claim 1, wherein the operations further comprise: providing, for presentation on a display device, a graphical user interface comprising a plurality of user interface elements, wherein each of the plurality of user interface elements corresponds to a different style which can be applied by the one or more machine-learned models for generating the output document having the particular document template; and receiving a selection of one or more of the plurality of user interface elements, wherein generating, via the one or more machine-learned models, the output document having the particular document template comprises applying one or more styles corresponding to the selection of the one or more of the plurality of user interface elements.

8. The computing device of claim 1, wherein the operations further comprise: providing, for presentation on a display device, a graphical user interface comprising a plurality of user interface elements, wherein each of the plurality of user interface elements corresponds to a different heading which can be applied by the one or more machine-learned models for generating sections of the output document having the particular document template; and receiving a selection of one or more of the plurality of user interface elements, wherein generating, via the one or more machine-learned models, the output document having the particular document template comprises applying one or more headings corresponding to the selection of the one or more of the plurality of user interface elements to sections of the output document.

9. The computing device of claim 1, further comprising: one or more databases configured to store a plurality of generative machine-learned models respectively associated with a plurality of different document types, and the operations further comprise retrieving, from among the plurality’ of generative machine-learned models, at least one generative machine-learned model associated with a document type indicated by common content of the plurality of source documents.

10. The computing device of claim 1, wherein the operations further comprise: receiving a plurality of training documents associated with a document type, wherein each of the plurality of training documents include one or more common features associated w ith the particular document template; andtraining the one or more machine-learned models to learn at least one of an intent, a style, or a format of the particular document template for the document type.

11. A computer-implemented method, comprising: receiving, by a computing system comprising one or more processors and one or more machine-learned models, a plurality of source documents; receiving, by the computing system, an input associated with a request to generate an output document based on the plurality of source documents and a particular document template; and generating, via the one or more machine-learned models, the output document having the particular document template.

12. The computer-implemented method of claim 11, further comprising: determining, by the one or more machine-learned models, at least one of a style, an intent, and a format associated with the output document, based on a content of each of the plurality of source documents.

13. The computer-implemented method of claim 11, further comprising: providing, for presentation on a display device of the computing system, a graphical user interface comprising a plurality of user interface elements, wherein each of the plurality of user interface elements corresponds to a sty le, an intent, or a format which can be applied by the one or more machine-learned models for generating the output document having the particular document template; and receiving a selection of one or more of the plurality of user interface elements, wherein generating, via the one or more machine-learned models, the output document having the particular document template comprises applying the style, the intent, or the format corresponding to the selection of the one or more of the plurality' of user interface elements.

14. The computer-implemented method of claim 11, further comprising: storing, in one or more databases, a plurality' of generative machine-learned models respectively associated with a plurality of different document types; andretrieving, from among the plurality of generative machine-learned models, at least one generative machine-learned model associated with a document type indicated by common content of the plurality of source documents.

15. The computer-implemented method of claim 11, further comprising: receiving a plurality of training documents associated with a document type, wherein each of the plurality of training documents include one or more common features associated with the particular document template; and training the one or more machine-learned models to learn at least one of an intent, a style, or a format of the particular document template for the document type.

16. A computing device for generating a document template, comprising: one or more memories configured to store instructions; and one or more processors configured to execute the instructions to perform operations, the operations comprising: receiving a selection of a plurality of sample documents; receiving an input to generate the document template based on the selection of the plurality7of sample documents; and in response to receiving the input, implementing one or more machine-learned models to generate the document template based on at least one of a style, a format, or an intent associated with the plurality of sample documents.

17. The computing device of claim 16, wherein the operations further comprise: receiving a plurality of source documents, receiving a further input associated with a request to generate an output document based on the plurality of source documents and the document template, and generating, via the one or more machine-learned models, the output document according to the document template.

18. The computing device of claim 17, wherein the operations further comprise: determining, by the one or more machine-learned models, the at least one of the sty le, the format, and the intent associated with the plurality of sample documents, based on a content and structure of each of the plurality of sample documents.

19. The computing device of claim 18, wherein the operations further comprise: providing, for presentation on a display device, a graphical user interface comprising a plurality of user interface elements, wherein at least one of the plurality of user interface elements is selectable to change the at least one of the style, the format, and the intent associated with the plurality of sample documents determined by the one or more machine- learned models; and receiving a selection of one or more of the plurality of user interface elements to change the at least one of the style, the format, and the intent associated with the plurality of sample documents determined by the one or more machine-learned models, wherein generating, via the one or more machine-learned models, the output document according to the document template comprises applying the at least one of the style, the format, and the intent changed according to the selection of the one or more of the plurality of user interface elements.

20. The computing device of claim 16, further comprising: one or more databases configured to store a plurality of generative machine-learned models respectively associated with a plurality of different document types, and the operations further comprise retrieving, from among the plurality of generative machine-learned models, at least one generative machine-learned model associated with a document type indicated by common content of the plurality of sample documents.