Generating content via a machine-learned model based on user-selected source content.

The computing system addresses inefficiencies in large language model interactions by automating content processing, enhancing efficiency and accuracy through direct user content analysis using machine-learned models.

DE202024107563U1Active Publication Date: 2025-05-08GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE202024107563
Authority / Receiving Office
DE · DE
Patent Type
Utility models
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-08
Estimated Expiration
2034-12-31

AI Technical Summary

Technical Problem

Current computing systems require significant user effort to process content with large language models, leading to inefficient resource usage due to switching between applications and windows.

Method used

A computing system configured to create summaries and identify key topics using machine-learned models directly from user-selected content, reducing the need for manual input and minimizing resource expenditure.

Benefits of technology

Enhances efficiency by automating content processing, saving resources and time, while improving the accuracy and reliability of information generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Computing device for generating content, comprising: one or more memories configured to store instructions; and one or more processors configured to execute instructions to perform operations, the operations including the following: Providing, in response to a selection of a multitude of content objects, a user interface that includes a first section and a second section, wherein the first section contains a summary description generated by one or more machine-learned models based on the multitude of content objects, and the second section contains a multitude of user interface elements configured to perform an operation with respect to at least one of the summary description or the multitude of content objects.
Need to check novelty before this filing date? Find Prior Art

Description

AREA

[0001] The disclosure generally relates to generating content via one or more machine-learned models based on source content selected (identified) by a user. For example, the disclosure relates to methods and computing devices for generating content by implementing a notebook application to obtain the source content to assist the user in managing content, organizing content, creating content, etc. BACKGROUND

[0002] With today's computing systems, large language models (LLMs) are capable of interacting with text content. For example, a user can copy content from a document and paste it into a chat box to send a query about the content to the LLM. The LLM can provide output (e.g., a summary) related to the content. SUMMARY

[0003] Aspects and advantages of embodiments of the disclosure will be set forth in part in the description below, or may be learned from the description, or may be learned by practice of the example embodiments.

[0004] In one or more example embodiments, a computing device is provided for generating, organizing, managing, and creating content.For example, the computing device includes: one or more memories configured to store instructions; and one or more processors configured to execute the instructions to perform operations, the operations comprising: providing, in response to a selection of a plurality of content objects, a user interface including a first portion and a second portion, the first portion including a summary description generated via one or more machine-learned models based on the plurality of content objects, and the second portion including a plurality of user interface elements configured to perform an operation with respect to at least one of the summary description or the plurality of content objects.

[0005] In some implementations, the operations further comprise: receiving a selection of the plurality of content objects; and implementing the one or more machine-learned models to generate the summary description based on the plurality of content objects.

[0006] In some implementations, the plurality of user interface elements includes a first user interface element comprising a suggested query, and the operations further comprise: receiving a selection of the first user interface element; and implementing the one or more machine-learned models to generate a response to the suggested query based on at least one of the summary description or the plurality of content objects.

[0007] In some implementations, the operations further comprise providing a third portion for the user interface to provide a dialog for displaying the suggested query and response.

[0008] In some implementations, the third section includes a quote user interface element that specifies a number of content objects from the plurality of content objects referenced by the one or more machine-learned models to generate the response.

[0009] In some implementations, the operations further comprise: in response to receiving a selection of the quote user interface element, providing, for display in a fourth portion of the user interface, content from one or more content objects of the plurality of content objects used to generate the response.

[0010] In some implementations, the third section includes a note creation user interface element, and the operations further include: in response to receiving a selection of the note creation user interface element, creating a note including content from the suggested query and the answer, and saving the note.

[0011] In some implementations, the plurality of user interface elements includes a first user interface element configured to generate a note, and the operations further comprise: providing a third portion to the user interface that includes content from one or more content objects of the plurality of content objects; receiving a selection of a portion of the content from the one or more content objects of the plurality of content objects; receiving a selection of the first user interface element; and in response to receiving the selection of the portion of the content and the first user interface element, generating, via the one or more machine-learned models based on the portion of the content, the note that includes a summary of the portion of the content.

[0012] In some implementations, the plurality of user interface elements includes a first user interface element configured to add content to an existing note, and the operations further comprise: providing a third portion to the user interface that includes content from one or more content objects of the plurality of content objects; receiving a selection of a portion of the content from the one or more content objects of the plurality of content objects; receiving a selection of the first user interface element; and in response to receiving the selection of the portion of the content and the first user interface element, adding the portion of the content to the existing note.

[0013] In some implementations, the first portion further includes at least one key topic user interface element comprising at least one key topic related to the summary description of the plurality of content objects, and the operations further include receiving a selection of the at least one key topic user interface element; and implementing the one or more machine-learned models to generate an output related to the at least one key topic based on at least one of the summary description or the plurality of content objects.

[0014] In some implementations, the operations further comprise providing a third portion for the user interface to provide a dialog for displaying the key topic and the output related to the at least one key topic.

[0015] In some implementations, the operations further comprise: providing a third portion for the user interface to provide for display at least one note generated via the one or more machine learned models based on the plurality of content objects.

[0016] In some implementations, the third section includes a quote user interface element that specifies a number of content objects from the plurality of content objects referenced by the one or more machine-learned models to generate the at least one note.

[0017] In some implementations, the operations further comprise: in response to receiving a selection of the quote user interface element, providing, for display in a fourth portion of the user interface, content from one or more content objects of the plurality of content objects used to create the note.

[0018] In some implementations, the operations further comprise: in response to receiving the selection of the quote user interface element, providing, for display in the fourth portion of the user interface, contextual content about the content from the one or more content objects of the plurality of content objects used to create the note.

[0019] In some implementations, the plurality of user interface elements includes a first user interface element configured to generate new content based on one or more notes, and the operations further comprise: providing a third portion for the user interface to provide for display a plurality of notes generated via the one or more machine learned models based on the plurality of content objects; receiving a selection of the plurality of notes; receiving a selection of the first user interface element; and in response to receiving the selection of the plurality of notes and the first user interface element, generating the new content based on the plurality of notes.

[0020] In some implementations, the operations further comprise: generating, via the one or more machine-learned models, a graphical image representing the plurality of content objects; and providing a folder containing the graphical image, the folder storing the plurality of content objects and a project file containing the summary description.

[0021] In some implementations, the second portion includes a text input field for receiving a query from a user, and the operations further include: implementing the one or more machine learned models to generate a response to the query based on at least one of the summary description or the plurality of content objects.

[0022] In one or more example embodiments, a computing device is provided for generating, organizing, managing, and creating content. For example, the computing device includes: one or more memories configured to store instructions; and one or more processors configured to execute the instructions to perform operations, the operations comprising: receiving input to create a notebook; receiving a selection of a plurality of content objects to add to the notebook;in response to receiving the selection of the plurality of content objects, implementing one or more machine-learned models to generate a summary description based on the plurality of content objects and at least one of a key topic user interface element indicating a topic of the plurality of content objects or a selectable user interface element indicating a query concerning the plurality of content objects; and providing a user interface including a first section and a second section, the first section including the summary description and the second section including the selectable user interface element.

[0023] In one or more example embodiments, a computer-implemented method for organizing, managing, and creating content is provided. The computer-implemented method includes providing, by a computing system and in response to a selection of a plurality of content objects, a user interface including a first portion and a second portion, the first portion including a summary description generated via one or more machine-learned models based on the plurality of content objects, and the second portion including a plurality of user interface elements configured to perform an operation with respect to at least one of the summary description or the plurality of content objects.

[0024] In one or more example embodiments, a computer-implemented method for organizing, managing, and creating content is provided. The computer-implemented method includes receiving, by a computing system, input to create a notebook; receiving, by the computing system, a selection of a plurality of content items to be added to the notebook; in response to receiving the selection of the plurality of content items, implementing, by the computing system, one or more machine-learned models to generate a summary description based on the plurality of content items and at least one of a key topic user interface element indicating a topic of the plurality of content items or a selectable user interface element indicating a query concerning the plurality of content items;and providing, by the computing system, a user interface including a first section and a second section, the first section including the summary description and the second section including the selectable user interface element;

[0025] In one or more example embodiments, a computer-readable medium (e.g., a non-transitory computer-readable medium) is provided that stores instructions executable by one or more processors of a computing system. In some implementations, the computer-readable medium stores instructions, which may include instructions to cause the one or more processors to perform one or more operations associated with one of the methods described herein (e.g., operations of the server computing system and / or operations of the computing device). The computer-readable medium may store additional instructions to perform other aspects of the server computing system and computing device and corresponding methods of operation as described herein.

[0026] These and other features, aspects, and advantages of various embodiments of the disclosure will be better understood by reference to the following description, drawings, and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the disclosure and, together with the description, serve to explain the principles involved. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] A detailed discussion of exemplary embodiments directed to a person of ordinary skill in the art is set forth in the description, which refers to the accompanying drawings, in which: The Fig. 1A-1B illustrate example systems according to one or more example embodiments of the disclosure; Fig. 2 illustrates a flow diagram of an exemplary, non-limiting computer-implemented method according to one or more example embodiments of the disclosure; Fig. 3 illustrates an example block diagram of a computing device according to one or more example embodiments of the disclosure; The Fig. 4A-4H illustrate example user interface screens of a notebook application according to one or more example embodiments of the disclosure; The Fig. 5A-5B illustrate example user interface screens of a notebook application according to one or more example embodiments of the disclosure; The Fig. 6A-6B illustrate further example user interface screens of a notebook application according to one or more example embodiments of the disclosure; Fig. 7 illustrates example notebooks or projects that may be represented in a particular manner according to one or more example embodiments of the disclosure; Fig. 8A illustrates a block diagram of an example computing system for organizing, managing, and creating content by implementing one or more machine-learned models with respect to user-selected source content, in accordance with one or more example embodiments of the disclosure; Fig. 8B illustrates a block diagram of an example computing device for organizing, managing, and creating content by implementing one or more machine-learned models with respect to user-selected source content, in accordance with one or more example embodiments of the disclosure; Fig. 8C illustrates a block diagram of an example computing device for organizing, managing, and creating content by implementing one or more machine-learned models with respect to user-selected source content, in accordance with one or more example embodiments of the disclosure. DETAILED DESCRIPTION

[0028] Reference will now be made to embodiments of the disclosure, one or more examples of which are illustrated in the drawings, wherein like reference numerals designate like elements. Each example is provided to illustrate the disclosure and is not intended to be limiting of the disclosure. Indeed, it will be apparent to those skilled in the art that various modifications and variations may be made to the disclosure without departing from the scope or spirit of the disclosure. For example, features illustrated or described in one embodiment may be utilized in another embodiment to provide yet another embodiment. Thus, the disclosure is intended to cover such modifications and variations that come within the scope of the appended claims and their equivalents.

[0029] The terms used in this specification are used to describe the exemplary embodiments and are not intended to limit and / or restrict the disclosure. The singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. In this disclosure, terms such as "including," "having," "comprising," and the like are used to define features, integers, steps, acts, elements, components, or combinations thereof, but do not preclude the presence or addition of one or more of the features, elements, steps, acts, elements, components, or combinations thereof.

[0030] It should be understood that although the terms first, second, third, etc. may be used throughout this specification to describe different elements, the elements are not limited by these terms. Instead, these terms are used to distinguish one element from another. For example, a first element may be referred to as a second element, and a second element may be referred to as a first element, without departing from the scope of the present disclosure.

[0031] The term "and / or" includes a combination of any of the plurality of listed items in question, or any one of the plurality of listed items in question. For example, the scope of the expression or phrase "A and / or B" includes the object "A," the object "B," and the combination of the objects "A and B."

[0032] Furthermore, the scope of the phrase “at least one of A or B” is intended to include all of the following: (1) at least one of A, (2) at least one of B, and (3) at least one of A and at least one of B. Likewise, the scope of the phrase “at least one of A, B, or C” is intended to include all of the following: (1) at least one of A, (2) at least one of B, (3) at least one of C, (4) at least one of A and at least one of B, (5) at least one of A and at least one of C, (6) at least one of B and at least one of C, and (7) at least one of A, at least one of B, and at least one of C.

[0033] According to current computing systems, large language models (LLMs) are capable of interacting with text content. However, current computing systems require significant effort to construct a specific request for processing by an LLM. For example, a user may need to copy and paste content from a document into a chat box to submit a query to the LLM about the content. This switching between multiple windows or applications results in significant unnecessary computational overhead and waste of resources (e.g., processor cycles).

[0034] According to examples of the disclosure, a computing system (computing platform, computing device) is configured to generate a new type of output (e.g., an outline, a report, a summary, etc.) via one or more machine-learned models based on source content provided to the computing system (e.g., by the user). For example, the computing system may be configured to receive source content selected by a user and generate, via one or more machine-learned models, a summary of the source content, including an identification of one or more topics related to the source content.

[0035] For example, a user may identify and select a subset of documents (e.g., four documents) from a large set of documents (a large document corpus) pertaining to a topic (e.g., modern American history in the 1990s), which is provided to the computational system. The computational system may include one or more machine-learned models configured to receive as input the selected documents and provide as output a summary (or report, article, outline, etc.) pertaining to the selected documents and an identification of key themes (e.g., via a document guide).

[0036] In some implementations, the computing system is configured to implement a semantic retrieval method (e.g., clustering) and one or more machine-learned models (e.g., one or more LLMs) to generate a summary, key themes, and suggested queries (e.g., questions) to produce a document guide for content identified or specified by the user (e.g., based on a body of text found in the content).

[0037] For example, in some implementations, the computing system may be configured to receive source content from the user. For example, the user may upload source content (e.g., documents, images, audio files, websites, videos, presentations, PDFs, etc.). In some implementations, the computing system may be configured, in response to the user uploading source content, to automatically generate information, including a summary of the source content, generate top themes found in the source content, generate suggested topics and questions to help the user further explore the source content, etc. The information may be presented via a user interface. The user interface may be configured to receive input from the user (e.g., via a touch input, a mouse click, etc.) on a user interface element corresponding to a theme, a question, etc.The computing system may be configured to respond to the input from the user in response to receiving the input, for example, by providing a response to the question or motivation query based on the source content via one or more machine-learned models.

[0038] In some implementations, the computing system may be configured to automatically generate information, including a report, outline, or rewrite of the original content, in response to the user uploading source content, to generate new content based on the source content identified (selected) by the user. For example, the user may request that the computing system identify a defined number of motifs from one or more documents, to summarize customer interactions that occur over a defined period of time (e.g., at least two weeks), to generate a defined number of ideas based on a source document, etc.

[0039] In some implementations, the source content that the LLMs use or reference may be selected (e.g., curated) by the user. For example, the user may indicate or believe that the selected source content is trustworthy (e.g., trusted source content, reliable source content, etc.) or has a higher priority than other content that does not have such a designation. Therefore, the one or more machine-learned models are configured to generate summaries of content or generate new content based on trusted source content, thereby improving the accuracy and reliability of information and data provided to the user.Further, the one or more machine-learned models are configured to answer questions about the source content based on the trusted source content, thereby improving the accuracy and reliability of information and data provided in response to questions posed by the user.

[0040] In some implementations, the computing system may be configured to discover, add, or remove source content. For example, the user may add or remove source content. For example, the user may provide input requesting that the computing system discover source content (e.g., by performing a search for scholarly articles on a certain topic), and the user may add the discovered source content as part of the selected source content that is considered trusted by the user (and / or the computing system).

[0041] In some implementations, the computing system may be configured to receive additional source content by the user creating a new note, by the user uploading the source content to the computing system, by adding the source content via a website, etc. The computing system may be configured to generate or receive metadata regarding the added source content. For example, the metadata may include one or more of a title, an author, a date of upload, a date associated with the creation of the source content, a uniform resource location (URL) associated with the source content, etc.

[0042] In some implementations, the computing system may be configured to delete or remove source content by the user selecting the source content and providing input requesting deletion of the source content (e.g., from the notepad application). In some implementations, the source content may be deleted as a source used in generating summaries, key themes, etc., in the notepad application, but an original copy of the source content may be retained elsewhere.

[0043] In some implementations, the computing system may be configured to receive user input via a text input field (e.g., an open-ended text input field). For example, the user input may be in the form of a question (e.g., "What did Nixon say about the use of motor vehicles in his speech?"). For example, the user input may be in the form of a motif or idea (e.g., "Nixon automobile crisis" or "What is this document about?").

[0044] The computing system may be configured to provide, via one or more machine-learned models, a response to the user input based on the selected source content. In some implementations, the computing system is configured to indicate the number of sources (citations) used to provide the response. In some implementations, the computing system is configured to provide a source (citation) for representation that was used for a particular location in the response. In some implementations, the computing system is configured to provide additional context regarding the source (citation) that was used for the particular location in the response. For example, the computing system may indicate the location (e.g.,a sentence or paragraph) from the source on which a section of the answer is based and further provide a preceding and / or following reference from the source to provide further context regarding the particular reference.

[0045] In some implementations, the computing system may be configured to store one or more locations (e.g., snippets) from a generated response (reply) to a query (question) entered by the user. For example, the one or more locations may be stored in a defined area of ​​a scratchpad application. The defined area may be referred to as a scratchpad, and each piece of information stored in the scratchpad may be referred to as a note. The one or more locations may be selected by the user for storage as the first note in the scratchpad. In some implementations, quotes may be stored as the second note in the scratchpad. In some implementations, the user may select (e.g., highlight) a specific location from a quote (source content) for storage in the scratchpad as the third note.In some implementations, the user can save his own position or comments as a written note (fourth note).

[0046] According to examples of the disclosure, the computing system may be configured to provide a scratchpad application configured to generate output (e.g., an outline, a report, a summary, etc.) via one or more machine-learned models based on source content provided to the scratchpad application (e.g., by the user). The notebook application may be configured to allow a user to create different projects to complete different tasks. Each project may be configured to function similarly to a folder through which a user can store various information about each project. In some implementations, a single scratchpad may correspond to or be dedicated to a particular project. In some implementations, the notebook application may be configured to receive the source content as defined by the user.The notebook application can be configured to add, delete, or modify projects based on input received from a user. Each project can be provided with a default name, a user-provided name, or a name generated by the notebook application (e.g., through one or more machine-learning models) based on the information stored in the project (e.g., based on the source content).

[0047] In some implementations, the notebook application may be configured to automatically generate (e.g., using one or more machine-learning models, one or more generative machine-learning models, etc.) a graphical image (e.g., an emoji, an icon, etc.) or a graphical animation corresponding to or representing the source content in response to the notebook application being provided with source content. In some implementations, the graphical image or graphical animation may be overlaid on a folder provided as a user interface element that, when selected, causes the folder to open and display the contents of the folder to the user. Additionally or alternatively, in some implementations, the notebook application may be configured to automatically generate (e.g., using one or more machine-learning models, one or more generative machine-learning models, etc.) a graphical image (e.g., an emoji, an icon, etc.) or a graphical animation corresponding to or representing the source content in response to the notebook application being provided with source content.Using one or more machine-learning models, one or more generative machine-learning models, etc., generate a text description (name) that corresponds to or represents the source content. The text description can be overlaid on the folder, provided as a user interface element that, when selected, causes the folder to open and display the folder's contents to the user.

[0048] One or more technical advantages of the disclosure include generating content via one or more machine-learned models based on specific content objects selected by a user. Current methods for generating output using a large language model (LLM) require a user to copy and paste content from a document into a chat field to submit a query to the LLM about the content. Switching between multiple windows or applications results in significant unnecessary computational time and waste of resources (e.g., processor cycles). In contrast to current methods, a notebook application and one or more machine-learned models can automatically generate a digest of the user-selected content objects (e.g., source content) in response to a user uploading content objects.Therefore, a user does not need to switch between applications or windows or provide a prompt.

[0049] Another technical advantage of the disclosure includes one or more machine-learned models providing suggested queries based on user-selected content objects, suggested key topics based on user-selected content objects, selectable chips based on the output of a response, and the like. A user may select a suggested query, and the one or more machine-learned models may be configured to provide a response to the query based on the user-selected content objects. Providing the suggested query automatically conserves computational resources (e.g., network resources including bandwidth, processor cycles, etc.) because the user is not required to enter the suggested query.

[0050] Another technical advantage of the disclosure includes one or more machine-learned models generating content based on user-selected content objects and / or user-selected notes. Content generation can save time and computational resources because it eliminates the need for a user to cut and paste content from multiple sources to generate new content (e.g., an outline, essay, report, etc.) based on multiple content objects.

[0051] Another technical advantage of the disclosure includes one or more machine-learned models generating a graphical image or animation for display overlaid on a folder to indicate content stored in the folder. The graphical image or animation can enhance search capabilities and save computing resources otherwise consumed by a user opening and closing folders that do not contain content the user is actually seeking.

[0052] With reference now to the drawings, Fig. 1A illustrates an exemplary system according to one or more exemplary embodiments of the disclosure. Fig. 1A illustrates an example of a system 1000 including a computing device 100, an external computing device 200, a server computing system 300, and external content 500 that may be in communication with each other via a network 400. For example, the computing device 100 and the external computing device 200 may include any of a personal computer, a smartphone, a tablet computer, a global location service device, a smartwatch, and the like. The network 400 may include any type of communications network, including a wired or wireless network, or a combination thereof.Network 400 may include a local area network (LAN), a wireless local area network (WLAN), a wide area network (WAN), a personal area network (PAN), a virtual private network (VPN), or the like. For example, wireless communication between elements of the exemplary embodiments may occur via a wireless LAN, Wi-Fi, Bluetooth, ZigBee, Wi-Fi Direct (WFD), Ultra Wideband (UWB), Infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), a radio frequency (RF) signal, and the like. For example, wired communication between elements of the exemplary embodiments may occur via a pair cable, a coaxial cable, a fiber optic cable, an Ethernet cable, and the like.Communication over the network 400 may occur using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), or protection schemes (e.g., VPN, Secure HTTP, SSL).

[0053] As explained in more detail below, in some implementations, the computing device 100 and / or the server computing system 300 may form part of an application system that may provide a tool for users to manage or organize information (e.g., documents, images, etc.) via, for example, one or more machine-learned models.

[0054] In some example embodiments, the server computing system 300 may obtain data from one or more of a source content data store 350, a user data store 360, and a machine-learned model data store 370 to implement various operations and aspects of the application system as disclosed herein. The source content data store 350, the user data store 360, and the machine-learned model data store 370 may be provided integrally with the server computing system 300 (e.g., as part of the one or more storage devices 320 of the server computing system 300) or may be provided separately (e.g., remotely). Furthermore, the source content data store 350, the user data store 360, and the machine-learned model data store 370 may be combined as a single data store (database) or may include a plurality of respective data stores. In a data store (e.g.,Data stored in one data store (e.g., the source content data store 350) may overlap with some data stored in another data store (e.g., the user data store 360). In some implementations, one data store (e.g., the machine-learned model data store 370) may reference data stored in another data store (e.g., the user data store 360).

[0055] In some examples, source content repository 350 may store any type of information or content. For example, source content repository 350 may include books, product manuals, legal opinions, academic papers, proprietary data files, patent documents, web pages, emails, forum posts, social media posts, videos, images, geographic information, or any other type or form of content that can be stored or accessed in digital form (e.g., in a database, storage device, etc.). In some implementations, information may be stored in source content repository 350 by the user selecting specific documents, images, or other content to be stored in source content repository 350.

[0056] In some examples, user data store 360 ​​may include information related to one or more user profiles, including a variety of user data, such as user preference data, user demographic data, user calendar data, user social media data, user travel history data, and the like. For example, user data store 360 ​​may include, among other things, email data, including text content, images, email-related calendar information, or contact information; social media data, including comments, ratings, check-ins, likes, invitations, contacts, or reservations; calendar application data, including dates, times, events, descriptions, or other content; virtual wallet data, including purchases, electronic tickets, coupons, or deals; scheduling data; location data; SMS data; or other suitable data associated with a user account.According to one or more examples of the disclosure, the data may be analyzed to determine user preferences regarding the creation, management, and / or organization of content, for example, to automatically generate a summary of a document in a particular manner or style, to automatically provide customized features related to content, to provide suggestions, recommendations, and / or questions related to a particular piece of content identified by the user as source content, etc.

[0057] User data store 360 ​​is provided to illustrate potential data that, in some embodiments, could be analyzed by computing device 100 and / or server computing system 300 to identify user preferences, provide recommendations, generate, manage, and / or organize content, etc. However, such user data may only be collected, used, or analyzed with the user's consent after being informed of what data is being collected and how it is being used. Furthermore, in some embodiments, the user may be provided with a tool (e.g., in a notebook application or via a user account) to revoke or change the scope of permissions.Furthermore, certain information or data may be processed in one or more ways before being stored or used, such that personally identifiable information is removed or stored in an encrypted manner. Thus, certain user information stored in user data storage 360 ​​may be accessible or inaccessible to computing device 100 and / or server computing system 300 based on permissions granted by the user, or such data may not be stored in user data storage 360 ​​at all.

[0058] The machine-learned model data store 370 may store machine-learned models that may be retrieved and implemented by the server computing system 300 to generate distilled or fine-tuned machine-learned models (e.g., distilled or fine-tuned generative machine-learned models), which in some implementations may also be provided to the computing device 100. The machine-learned model data store 370 may also store distilled or fine-tuned machine-learned models (e.g., distilled or fine-tuned generative machine-learned models) that may be retrieved and implemented by the computing device 100. In some implementations, the computing device 100 may retrieve and implement machine-learned models that are large-parameter models that have not been fine-tuned or distilled.The machine-learned models (including large parameter models and distilled or fine-tuned models) stored in the machine-learned model data store 370 may include generative machine-learned models, each associated with different types of content (e.g., different genres or topics, different types of content, including images, videos, and text, different varieties of content, including outlines, reports, spreadsheets, etc.). The machine-learned models may include large language models (e.g., the Bidirectional Encoder Representations from Transformers (BERT) large language model) and general, multimodal models (e.g., Gemini). The machine-learned models may include generative artificial intelligence (AI) models (e.g.,Bard) that can implement Generative Adversarial Networks (GAN), Transformers, Variational Autoencoders (VAE), Neural Radiance Fields (NeRF), and the like.

[0059] External content 500 may be any form of external content, including news articles, web pages, video files, audio files, written descriptions, reviews, game content, social media content, photographs, commercial offers, transportation methods, weather conditions, sensor data obtained by various sensors, or other suitable external content. Computing device 100, external computing device 200, and server computing system 300 may access external content 500 via network 400. External content 500 may be searched by computing device 100, external computing device 200, and server computing system 300 using known search methods, and search results may be ranked by relevance, popularity, or other suitable attributes, including location-specific filtering or advertising.

[0060] With reference now to Fig. 1B, exemplary block diagrams of a computing device and a server computing system will now be described in accordance with one or more exemplary embodiments of the disclosure. Although the computing device 100 in Fig. 1B, features of the computing device 100 described in this document are also applicable to the external computing device 200.

[0061] Computing device 100 may include one or more processors 110, one or more storage devices 120, an application system 130, a positioning device 140, an input device 150, a display device 160, an output device 170, and a capture device 180. Server computing system 300 may include one or more processors 310, one or more storage devices 320, and an application system 330.

[0062] For example, the one or more processors 110, 310 may be any suitable processing device that may be included in a computing device 100 or a server computing system 300. For example, the one or more processors 110, 310 may include one or more of a processor, processor cores, a controller, and an arithmetic logic unit, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an image processor, a microcomputer, a field-programmable array, a programmable logic unit, an application-specific integrated circuit (ASIC), a microprocessor, a microcontroller, etc., and combinations thereof, including any other device capable of responding to and executing instructions in a defined manner.The one or more processors 110, 310 may be a single processor or a plurality of processors that are operatively connected, for example, connected in parallel.

[0063] The one or more storage devices 120, 320 may include one or more non-transitory computer-readable storage media, including read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and flash memory, a USB drive, a volatile storage device, including random access memory (RAM), a hard drive, floppy disks, a Blu-ray disc, or optical media such as CD-ROM discs and DVDs, and combinations thereof. However, examples of the one or more storage devices 120, 320 are not limited to the above description, and the one or more storage devices 120, 320 may be implemented by other various devices and structures, as would be understood by one of ordinary skill in the art.

[0064] For example, the one or more storage devices 120 may also include data 122 and instructions 124 that may be accessed, manipulated, created, or stored by the one or more processors 110.In some example embodiments, such data may be accessed and used as input to implement the notebook application 132 and execute the instructions to perform operations, including: providing a user interface including a first portion and a second portion, the first portion including a text summary generated via one or more machine-learned models based on a plurality of documents selected by a user, and the second portion including a plurality of user interface elements to perform an operation on the text summary, as described in accordance with examples of the disclosure.

[0065] For example, the one or more storage devices 320 may also include data 322 and instructions 324 that may be accessed, manipulated, created, or stored by the one or more processors 310.In some example embodiments, such data may be accessed and used as input to implement the notebook application 332 and execute the instructions to perform operations, including: providing a user interface including a first portion and a second portion, the first portion including a text summary generated via one or more machine-learned models based on a plurality of documents selected by a user, and the second portion including a plurality of user interface elements to perform an operation on the text summary, as described in accordance with examples of the disclosure.

[0066] In some example embodiments, computing device 100 includes an application system 130. For example, application system 130 may include notebook application 132 and a document application 134 (e.g., a word processing application, a spreadsheet application, a presentation application, an image application, etc.). Application system 130 may include various other applications, including text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, mapping applications, social media applications, navigation applications, etc.

[0067] According to examples of the disclosure, the notebook application 132 may be executed by the computing device 100 to provide a way for a user of the computing device 100 to organize, manage, create, and interact with content, particularly content curated or selected by the user. In some implementations, the notebook application 132 may be part of the document application 134 or a standalone application. The notebook application 132 may be configured to be dynamically interactive according to various user inputs. Example implementations of the notebook application 132 are described herein, but the disclosure is not limited to these examples, as various modifications may be made to the embodiments described herein.

[0068] In some examples, one or more aspects of notebook application 132 may be implemented by notebook application 332 of server computing system 300, which may be remotely located to organize, manage, create, and interact with content in response to receiving input from a user. In some examples, one or more aspects of notebook application 332 may be implemented by notebook application 132 of computing device 100 to organize, manage, create, and interact with content in response to receiving input from a user.

[0069] According to examples of the disclosure, the document application 134 may be executed by the computing device 100 to provide a way for a user of the computing device 100 to organize, manage, create, and interact with content, particularly content curated or selected by the user. The document application 134 may be any type of application related to documents (e.g., in a text or image format) and may include word processing applications, spreadsheet applications, presentation applications, visual applications, Portable Document Format file applications, etc. In some implementations, the notebook application 132 and the document application 134 may interact with each other. For example, content from a document created via the document application 134 may be uploaded or saved for use with the notebook application 132.In some implementations, the notebook application 132 may be configured to generate a document (e.g., a report, an outline, a presentation, a spreadsheet) that may be compatible with (opened by or exported to) the document application 134.

[0070] In some examples, document application 134 may be a dedicated application specifically configured to provide a particular service. In other examples, document application 134 may be a general-purpose application (e.g., a web browser) and may provide access to a variety of different services over network 400.

[0071] In some example embodiments, computing device 100 includes a positioning device 140. Positioning device 140 may determine a current geographic location of computing device 100 and communicate such a geographic location to server computing system 300 via network 400. Positioning device 140 may be any device or circuit for analyzing the position of computing device 100. For example, positioning device 140 may determine the actual or relative position using a satellite navigation positioning system (e.g.,a GPS system, a Galileo positioning system, the Global Navigation Satellite System (GLONASS), the BeiDou Satellite Navigation and Positioning System), an inertial navigation system, a dead reckoning system based on an IP address using triangulation and / or proximity to cell towers or Wi-Fi hotspots, and / or other suitable techniques for determining a position of the computing device 100.

[0072] Computing device 100 may include an input device 150 configured to receive input from a user, and may, for example, be one or more of a keyboard (e.g., a physical keyboard, virtual keyboard, etc.), a mouse, a joystick, a button, a switch, an electronic pen or stylus, a gesture recognition sensor (e.g., to detect a user's gestures, including movements of a body part), an input sound device, or a voice recognition sensor (e.g., a microphone to receive voice input, such as a voice command or query), a trackball, a remote control, a mobile phone (e.g., a cell phone or smartphone), a tablet PC, a pedal or footswitch, a virtual reality device, and so on. Input device 150 may also be implemented, for example, by a touch-sensitive display having touchscreen capability.For example, the input device 150 may be configured to receive input from a user associated with the input device 150 to select content to be organized or managed, to select queries or actions with respect to content curated or selected by the user, etc.

[0073] Computing device 100 may include a display device 160 that displays user-visible information (e.g., a user interface screen). For example, display device 160 may be a non-contact sensitive display or a touch-sensitive display. Display device 160 may include, for example, a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, an active matrix organic light-emitting diode (AMOLED) display, a flexible display, a 3D display, a plasma display panel (PDP), a cathode ray tube display (CRT), and the like. However, the disclosure is not limited to these example displays and may include other types of displays.The display device 160 may be used by the application system 130 provided on the computing device 100 to display to a user information relating to an input (e.g., information relating to a document, a note, a project, etc., a user interface screen having user interface elements selectable by the user, etc.).

[0074] The computing device 100 may include an output device 170 to provide output to the user and may include, for example, one or more of an audio device (e.g., one or more speakers), a haptic device to provide haptic feedback to a user (e.g., a vibrator), a light source (e.g., one or more light sources such as LEDs that provide visual feedback to a user), a thermal feedback system, and the like.

[0075] Computing device 100 may include a capture device 180 capable of capturing media content according to various examples of the disclosure. For example, capture device 180 may include an image capture device 182 (e.g., a camera) configured to capture images (e.g., photos, video, and the like). For example, capture device 180 may include a sound capture device 184 (e.g., a microphone) configured to capture sound or audio (e.g., an audio recording) from a location. The media content captured by capture device 180 may be transmitted, for example, over network 400 to one or more of server computing system 300, source content data store 350, user data store 360, and machine-learned model data store 370.For example, in some implementations, media content captured by capture device 180 may be selected by a user as source content for use in creating a note related to a project. For example, the media content may be provided as input to one or more machine-learned models to create a note.

[0076] According to exemplary embodiments of the disclosure, server computing system 300 may include one or more processors 310 and one or more memory devices 320, as described herein. Server computing system 300 may also include an application system 330 similar to application system 130 described herein.

[0077] For example, the application system 330 may include a notebook application 332 that performs similar functions as discussed above with respect to the notebook application 132. In some implementations, one or more machine-learned models (e.g., generative machine-learned models, large language models, etc.) associated with the application system 330 may be configured to organize, manage, create, and interact with content based on source content curated or selected by a user. For example, one or more machine-learned models (e.g., generative machine-learned models, large language models, etc.) associated with the application system 330 may be configured to perform a first action (e.g.,generate a summary or document guide with respect to source content selected by a user), while computing device 100 may be configured to perform a second action (e.g., generate suggested actions, generate an outline or learning guide based on a plurality of notes stored in a scratchpad). For example, a particular action to be performed by application system 330 may vary according to a network status (e.g., available bandwidth, channel utilization status, latency status, throughput rate, etc.). In some implementations, one or more machine-learned models associated with application system 330 may be configured to process user input to generate information (e.g., semantic information) that is then provided to one or more other machine-learned models (e.g.,generative machine learned models, large language models, etc.) associated with the application system 330 may be provided as input to generate the content to be deployed to the notebook application 132 and / or the notebook application 332 with respect to a project.

[0078] Examples of the disclosure also relate to computer-implemented methods for providing a user interface for organizing, managing, and creating content by implementing one or more machine-learned models with respect to source content selected by a user. Fig. 2 illustrates a flow diagram of an exemplary, non-limiting computer-implemented method according to one or more example embodiments of the disclosure. Fig. 3 illustrates a block diagram of a notebook application according to one or more example embodiments of the disclosure.

[0079] The flow chart from Fig. 2 illustrates a method 2000 for providing a user interface for organizing, managing, and creating content by implementing one or more machine-learned models with respect to user-selected source content. Although shown in a particular sequence or order, the order of the processes may be changed unless otherwise noted. Thus, the illustrated embodiments are only examples, and the illustrated processes may be performed in a different order, and some processes may be performed in parallel. Additionally, in various embodiments, one or more processes may be omitted. Thus, not all processes are required in every embodiment. Other process flows are possible.

[0080] With reference to Fig. 2, at operation 2100, the method 2000 includes a computing device receiving input from a user regarding the selection of source content. As described herein, the computing device may be implemented as computing device 100, server computing system 300, or combinations thereof. For example, the input may be provided by the user via input device 150. For example, the input may be provided by selecting certain files or documents to be uploaded to the computing device for use by notebook application 132. In some implementations, the content may be uploaded from local storage, from another application (e.g., a portable document file application), from copied text, or from a website. The selected files, text, documents, etc., may be referred to as source content.In some implementations, the source content may be a subset of a larger body of content. For example, the input may be provided to or input into notebook application 132 or notebook application 332.

[0081] In some implementations, a response to the input selecting the source content may be processed on computing device 100 without involving server computing system 300. In some implementations, the input selecting the source content may be transmitted from computing device 100 to server computing system 300, and at least a portion of the response to the input may be processed by server computing system 300. For example, the input regarding the selection of the source content may be provided on computing device 100, and server computing system 300 may be configured to perform an operation in response to receiving an indication of the input.

[0082] The computing device may be configured, at operation 2200, to implement one or more machine-learned models with respect to the selected source content to generate a document guide. In some implementations, the document guide (source guide) generated by the one or more machine-learned models may include a summary of the source content and important topics related to the source content. In some implementations, the document guide may further include one or more suggested queries (e.g., questions) that may be provided as a selectable user interface element.

[0083] For example, the computing device may obtain information indicating that the user has selected source content. The computing device may process the source content with one or more machine-learned models (e.g., one or more large-scale language models) to obtain a speech output. The computing device may then use the one or more machine-learned models (e.g., one or more large-scale language models) to generate a summary output. In particular, a machine-learned large-scale language model may be trained to process a variety of outputs to generate a speech output.For example, the machine-learned large language model may process an embedding generated by a machine-learned embedding generation model, portions of source content identified using the embedding generation model, speech outputs generated using the machine-learned large language model or another model, etc.

[0084] The computing device may be configured to receive input to perform an action with respect to the document guide at operation 2300. The computing device may be configured to perform the action in response to receiving the input at operation 2400. For example, the input may be selecting a suggested query, and the action may include providing an answer to the question by implementing the one or more machine-learned models with respect to the source content. For example, the input may be a text input that asks a question, and the action may include providing an answer to the question by implementing the one or more machine-learned models with respect to the source content.For example, the input may be a selection of a section of the summary, and the action may include providing output that indicates specific sources from the source content that were used to generate the text associated with the selection of the section of the summary.

[0085] With reference to Fig. 3, the notebook application 3100 (which may correspond to the notebook application 132 and / or the notebook application 332) may include a conditioning parameter generator 3110, one or more sequence processing models 3120, one or more large language models 3130, and one or more generative machine-learned models 3140. The notebook application 3100 may receive input 3200 from a user, as described above with respect to operation 2100 and operation 2300 of Fig. 2. The conditioning parameter generator 3110 may be configured to generate conditioning parameters based at least in part on the input, wherein the conditioning parameters provide values ​​for one or more conditions associated with generated content that relates at least in part to the input 3200 and the source content 3400 selected by the user.

[0086] For example, source content 3400 may include documents of any type (e.g., in digital form) and may include books, product manuals, legal opinions, academic papers, proprietary data files, patent documents, web pages, emails, forum posts, social media posts, videos, images, geographic information, or any other type or form of content that can be stored or accessed in digital form (e.g., in a database, storage device, etc.). In some implementations, source content 3400 may be stored in source content repository 350 by the user selecting particular documents, images, or other content to be stored in source content repository 350. In some implementations, source content 3400 may be stored on computing device 100 or server computing system 300.

[0087] To generate the conditioning parameters, the conditioning parameter generator 3110 may be configured to retrieve values ​​for the one or more conditions associated with the input. To generate the conditioning parameters, the conditioning parameter generator 3110 may, for example, be configured to extract the values ​​for the one or more conditions from the input. The input may include information indicating the user's intent or requirements.In some implementations, the conditioning parameter generator 3110 (or the one or more sequence processing models 3120 or the one or more large language models 3130) may be configured to extract information from the input 3200 to identify values ​​for the one or more conditions, and the conditioning parameter generator 3110 may be configured to generate the conditioning parameters based on the extracted values. For example, the input itself may identify a color to be used for headings in a generated document (e.g., "blue font for the title") or an attribute or feature (e.g., "dotted bullets") that can be used to generate the conditioning parameters for generating a document pertaining to the source content.

[0088] To generate the conditioning parameters, the conditioning parameter generator 3110 may be configured to derive the values ​​for the one or more conditions from the input. The input may include information indicating the user's intent or requirements. In some implementations, the conditioning parameter generator 3110 (or the one or more sequence processing models 3120 or the one or more large language models 3130) may be configured to derive information from the input 3200 to identify values ​​for the one or more conditions, and the conditioning parameter generator 3110 may be configured to generate the conditioning parameters based on the derived values. For example, the input may include a reference to a length ("short," "long," etc.).) of the summary to be generated based on the source content or other document to be generated, and the conditioning parameter generator 3110 (or the one or more sequence processing models 3120 or the one or more large language models 3130) may be configured to derive a value based on the input. For example, an input requesting the notebook application 3000 to generate a "short" essay may derive a value of approximately 500 words, while a "long" essay may be associated with a value of approximately 2000 words. For example, the notebook application 3100 may be configured to determine a derived value based on information about external content 3300.

[0089] In some implementations, the conditioning parameter generator 3110 may be configured to derive the values ​​for the one or more conditions from the input by providing the input to one or more sequence processing models 3120, wherein the one or more sequence processing models 3120 are configured to output the values ​​for the one or more conditions in response to or based on the query. The one or more sequence processing models 3120 may include one or more machine-learned models configured to process and analyze sequential data and to process data that occurs in a particular order or sequence, including time series data, natural language text, or other data with a temporal or sequential structure.

[0090] The one or more sequence processing models 3120 may receive input including text and tokenize the input by splitting the text sequence into small units (tokens) to provide a structured representation of the input sequence. The one or more sequence processing models 3120 may represent the tokens as vectors in a continuous vector space by mapping each token to a high-dimensional vector, where the relationships between tokens (words) are reflected in the geometric relationships between their corresponding vectors. For example, the one or more sequence processing models 3120 may receive input including the text "How did the Cold War end?" and tokenize the input by splitting the text sequence into small units (tokens) (e.g., "How," "Cold War," and "ended"), thereby providing a structured representation of the input sequence.In a word embedding, semantically similar words are closer together in the vector space. For example, the vectors for "war" and "battle" might be close to each other due to their semantic relationship, while the vectors for "war" and "peace" might be far apart compared to the vectors for "war" and "battle."

[0091] The one or more large-scale language models 3130 may be or otherwise include a model trained with a large corpus of language training data in a manner that provides the one or more large-scale language models 3130 with the ability to perform multiple language tasks. For example, the one or more large-scale language models 3130 may be trained to perform summarization tasks, conversation tasks, simplification tasks, opposing viewpoint tasks, etc. In particular, the one or more large-scale language models 3130 may be trained to process a variety of outputs to generate a language output. For example, the one or more large-scale language models 3130 may process an embedding generated by a machine-learned embedding generation model, portions of source content (e.g.,a block(s) of the document) identified using an embedding generation model, process speech outputs generated using the one or more large language models 3130 or another model, etc.

[0092] The one or more generative machine-learned models 3140 may include a deep neural network or a generative adversarial network (GAN), variational autoencoders, machine-learned stable diffusion models, visual transformers, neural radiance fields (NeRF), etc., to generate content (e.g., a summary, response to a query, etc.) with values ​​for conditions associated with one or more features. For example, the computing device may include a database (e.g., the machine-learned model data store 370) configured to store a plurality of generative machine-learned models, each associated with a plurality of different types of content (e.g., different genres or topics, different types of content, including images, videos, and text, different varieties of content, including outlines, reports, spreadsheets, etc.).In some implementations, the computing device may be configured to retrieve from the one or more generative machine-learned models 3140 a generative machine-learned model associated with a particular type of content concerning the input.

[0093] In some implementations, the one or more generative machine-learned models 3140 may be trained with a large content dataset (e.g., a large corpus of language training data) with corresponding information about the conditions associated with the content. During training, the one or more generative machine-learned models 3140 learn relationships between elements in an output (e.g., content) and conditions that affect it. This may include the computing device adjusting the internal parameters of each generative machine-learned model to generate realistic or accurate content (e.g., grammatically correct content, coherent content, etc.) based on the training data. The one or more generative machine-learned models 3140 may be trained with one or more training datasets, including a plurality of reference images of the location.The one or more training data sets may contain values ​​for the one or more conditions.

[0094] In some implementations, the one or more generative machine-learned models 3140 are configured to make decisions to generate content based on the conditioning parameters (and corresponding values ​​for the one or more conditions), to generate the document guide 3500 in response to receiving the selection of the source content 3400 and / or to generate responsive content 3600 corresponding to the content generated in response to the input to perform an action with respect to the document guide, etc.

[0095] In some implementations, server computing system 300 may provide (transfer) content or a portion of the generated content to computing device 100, or server computing system 300 may provide access to the generated content to computing device 100. For example, document guide 3500 may be generated at server computing system 300 and stored on one or more computing devices (e.g., one or more of computing device 100, external computing device 200, server computing system 300, external content 500, source content data store 350, user data store 360, etc.).

[0096] In some implementations, after a document guide is generated and / or after an action is performed with respect to the document guide, the user may provide feedback or further input regarding the content generated based on the provided source content and / or a query provided via the user, and repeat one or more of operations 2100 to 2400.

[0097] Examples of the disclosure are also directed to user-facing aspects by which a user can manage content, organize content, create content, etc., via a notebook application configured to implement one or more machine-learned models with respect to source content selected by the user. For example, the Fig. 4A to 4H illustrate examples of actions that may be implemented for a project in which a document guide is generated via one or more machine-learned models based on source content selected by a user, in accordance with one or more example embodiments of the disclosure.

[0098] For example, Fig. 4A illustrates a first user interface screen (e.g., a home user interface screen, a home graphical user interface, etc.) of a notebook application according to one or more example embodiments of the disclosure.

[0099] In Fig. 4A, the first user interface screen 4100 illustrates a user interface (e.g., an introductory screen) that provides information about the notebook application 3100. Specifically, the notebook application 3100 is configured to present for display the first user interface screen 4100, which includes various information 4110 regarding features available in the notebook application 3100.

[0100] As in Fig. As illustrated in Figure 4B, the notebook application 3100 is further configured to present for display a second user interface screen 4200 including a first user interface element 4210. For example, the first user interface element 4210 is associated with enabling a user to create a new notebook (a new project) with which a user can manage content, organize content, create content, etc., based on source content that the user can select or curate.

[0101] As in Fig. As illustrated in Figure 4C, the notebook application 3100 is further configured, in response to a user providing input to create a new notebook (e.g., via selection of the first user interface element 4210), to present a third user interface screen 4300 for display. The third user interface screen 4300 includes a first portion 4310 having a plurality of selectable user interface elements corresponding to locations from which the source content may be uploaded. For example, the first user interface element 4312 corresponds to a storage location, which may be associated with a local computing device or a remote server system (e.g., a cloud server) or other storage device (e.g., a portable storage device).For example, the second user interface element 4314 corresponds to a Portable Document Format file, the third user interface element 4316 corresponds to copied text, and the fourth user interface element 4318 corresponds to content that can be uploaded from a specific website or URL.

[0102] As in Fig. 4D, the notebook application 3100 is further configured to, in response to the selection of one of the plurality of selectable user interface elements corresponding to locations from which the source content may be uploaded, as described with respect to Fig. 4C, presenting a fourth user interface screen 4400 for display. The fourth user interface screen 4400 includes a first portion 4410 having a plurality of selectable source content objects (e.g., a plurality of documents, images, videos, etc.). For example, the first user interface element 4412 corresponds to a first selected document, the second user interface element 4414 corresponds to a second selected document (e.g., a Portable Document Format file), and the third user interface element 4416 corresponds to a third selected document. Fig. 4D illustrates that the user can curate or select specific source content objects that can be used to create a notebook or project and that can be used by one or more machine-learning models as input data to organize content, manage content, create content, etc.

[0103] As in Fig. 4E, the notebook application 3100 is further configured to, in response to selecting one or more content objects from the plurality of source content objects related to Fig. 4D, present a fifth user interface screen 4500 for display. The fifth user interface screen 4500 includes a first section 4510, a second section 4520, a third section 4530, and a fourth section 4540. Each section of the fifth user interface screen 4500 may correspond to a sub-area or field of the fourth user interface screen 4400 and may be associated with different functionality.

[0104] For example, the first section 4510 corresponds to a document guide (also referred to as a resource guide) that includes a summary section 4512 and a key topic section 4514. The notebook application 3100 may be configured to generate the content (e.g., a text description) associated with the summary section 4512 by using one or more machine-learned models, as described herein with respect to Fig. 3, based on the selected source content (e.g. as described in relation to Fig. 4D). For example, the summary sub-area 4512 may provide a brief summary associated with one or more of the content objects that comprise the selected source content. Similarly, the notebook application 3100 may be configured to generate the content (e.g., a text description) associated with the key topic sub-area 4514 by implementing one or more machine-learned models, as described herein with respect to Fig. 3, based on the selected source content (e.g. as described in relation to Fig. 4D). For example, the key topic sub-area 4514 may include one or more user interface elements that identify themes or important themes associated with one or more of the content objects comprising the selected source content. Further, the notebook application 3100 may be configured to generate output in response to a selection of one of the user interface elements in the key topic sub-area 4514. The output may be, for example, a text summary or a text explanation regarding the key topic corresponding to the selected user interface element.The output may be provided on a separate user interface screen or provided in a different portion of the fifth user interface screen, which the notebook application 3100 is configured to generate in response to the selection of one of the user interface elements in the key topic sub-area 4514.

[0105] For example, the second section 4520 corresponds to a source content section (e.g., a context window) that includes information 4522 from at least a portion of a content object from the source content. The notebook application 3100 may be configured to reproduce at least a portion of a content object from the source content in the second section 4520. In some implementations, the content in the source content section may correspond to a portion of a content object used to generate the summary section 4512.

[0106] For example, the third portion 4530 corresponds to a notes sub-area (e.g., a scratchpad) that may include one or more notes that may be generated via various methods as described herein (e.g., automatically generated by the notebook application 3100, manually entered by a user, automatically generated by the notebook application 3100 in response to the selection of a user interface element corresponding to an action to be performed, etc.).

[0107] For example, the fourth section 4540 corresponds to a query sub-area, which may include one or more user interface elements for sending or providing a query to the notebook application 3100 regarding the source content. For example, the fourth section 4540 includes a plurality of user interface elements 4542 corresponding to suggested questions or actions related to the source content. For example, the notebook application 3100 may be configured to generate the suggested questions or actions based on information included in the source content. For example, the notebook application 3100 may be configured to additionally generate the suggested questions or actions based on dialogue history (e.g., previous questions or queries), user data (e.g., user preferences, user attributes, etc.), and other contextual information.The fourth section 4540 may further include a text input field 4544 that allows a user to provide input (e.g., via a keyboard, via voice input, etc.) to query the notebook application 3100. The fourth section 4540 may further include a user interface element 4546 that indicates the number of content objects that comprise the source content. For example, in . Fig. 4E indicates that the notebook application 3100 used three sources to generate the summary subsection 4512.

[0108] With reference to Fig. Figure 4F illustrates an example user interface screen showing an input question and an output response related to the source content. For example, Fig. 4F, the notebook application 3100 is further configured to, in response to receiving a query (e.g., a text query entered via the text input field 4544 from Fig. 4E) to present a sixth user interface screen 4600 for display. For example, the sixth user interface screen 4600 includes a first section 4610 corresponding to a dialog sub-area, a second section 4620 corresponding to a source sub-area, and a third section 430 corresponding to a notes sub-area (e.g., a scratchpad).

[0109] For example, the first section 4610 includes a prompt area 4612 corresponding to the text query and an answer area 4614 corresponding to the answer to the text query. In some implementations, the notebook application 3100 is configured to generate the answer by implementing one or more machine-learned models in response to receiving the text query as input and with respect to the source content 3400. For example, if a user enters a question (e.g., "How did the Cold War influence American foreign policy?") via the text input field 4544, as with respect to Fig. 4E, the notebook application 3100 may be configured to provide the sixth user interface screen 4600 and generate a response as indicated in the response area 4614. As indicated in the response area 4614, the number of references (source content objects) that the one or more machine-learned models utilized to generate the response may be indicated by a first user interface element 4616. In the example of Fig. 4F, three references were used to generate the response. The response area 4614 further includes a selectable second user interface element 4618 that, when selected, causes the response to be saved as a note in the third section 4630 corresponding to the notes sub-area (e.g., a scratchpad), which may include one or more notes generated via various methods as described herein (e.g., automatically generated by the notebook application 3100 in response to the selection of the second user interface element 4618, etc.).

[0110] In some implementations, one or more sections of the response area may include information that is selectable and that, when selected, may cause additional information related to the selected information to be displayed. For example, in Fig. 4F, the text "Containment Policy" may be highlighted, bold, underlined, or visually distinguished to indicate that the text is selectable (e.g., a clickable chip) and that additional information related to the text is available. The notebook application 3100 may be configured to provide the additional information (e.g., by implementing one or more machine-learned models based on the selected source content) to provide additional information related to the text in response to the selection of the text.

[0111] The second section 4620 may correspond to a source sub-area and include the content objects 4622 that comprise the source content. In some implementations, the content objects 4622 may correspond to content objects that utilize the one or more machine-learned models to generate the response. In some implementations, the notebook application 3100 may be configured to dynamically modify or regenerate a response in the response area 4614 in response to receiving, via the user interface element 4624, an additional content object to be added as source content. Additionally or alternatively, in some implementations, the notebook application 3100 may be configured to dynamically modify or regenerate a response in the response area 4614 in response to receiving, via the user interface element 4626, a deselection of a content object from the list of content objects in the second section 4620 (e.g.,by deactivating the check box for one or more of the content objects in the second section 4620) to dynamically change or regenerate a response in the response area 4614.

[0112] With reference to Fig. 4G, an example user interface screen includes an example notes section (scratchpad) for a project according to examples of the disclosure. For example, in Fig. 4G, the notebook application 3100 is further configured to, in response to receiving a selection of the second user interface element 4618 (e.g., as shown in Fig. 4F), which, when selected, causes the response to be stored as a note 4712 in the third section 4710 corresponding to the notes sub-area (e.g., a scratchpad), which may include one or more notes generated via various methods as described herein (e.g., automatically generated by the notebook application 3100 in response to the selection of the second user interface element 4618, etc.), to present a seventh user interface screen 4700 for display. In Fig. 4G, user interface element 4714 indicates the number of content objects that one or more machine-learned models utilized to generate the response for note 4712. Furthermore, user interface element 4714 may be configured to be selectable such that, in response to user interface element 4714 being selected, a list of the content objects (citations) from the source content used to generate the response may be provided for display.

[0113] With reference to Fig. 4H, an example user interface screen includes a notes section (scratchpad) for a project according to examples of the disclosure. For example, the notebook application 3100 is Fig. 4H is further configured to present an eighth user interface screen 4800 for display in response to receiving a selection of a content object 4816 from a list 4814 of content objects (quotes) from the source content used by the one or more machine-learned models to generate the response stored in the note 4812 provided for display in the first section 4810. The eighth user interface screen 4800 further includes a second section 4820 corresponding to a source sub-area. The notebook application 3100 is in Fig. 4H configured to provide, in response to receiving the selection of the content object 4816 from the list 4814 of content objects (quotes) from the source content used by the one or more machine-learned models to generate the response stored in the note 4812, information 4822 relating to the selected object 4816 for display in the second section 4820.

[0114] In some implementations, the information 4822 may include information from the content object 4816 used to generate the response. For example, the notebook application 3100 may be configured to reference metadata associated with the response to link back to the information 4822. The metadata may specify a location of information from a content object used to generate the response. Further, the information 4822 may correspond to or include a particular location from the content object used to generate the response. For example, the notebook application 3100 may be configured to cause the particular location to be displayed in a visually distinctive manner in the second section 4820 (e.g., highlighted, bold, increased font size, underlined, italicized, etc.).For example, the notebook application 3100 may be configured to cause additional locations that appear before and / or after the specific location to be displayed in the second section 4820. This additional information may provide further context for the user regarding the information used to generate the answer. For example, the notebook application 3100 may be configured to highlight specific content objects used to generate the answer in the note 4812, as well as specific locations from the specific content objects used to generate the answer in the note 4812. Thus, a user can easily and visually identify where evidence for an answer can be found within a content object.

[0115] In some implementations, the information 4822 from the selected content object 4816 used to generate the response may be truncated or shown in its entirety. For example, if the information 4822 falls below a threshold, all text from the selected content object 4816 may be shown in the second portion 4820 and used by the one or more machine-learned models to generate a response (e.g., to a text query). For example, if the information 4822 exceeds the threshold, the notebook application 3100 may be configured to implement a semantic retrieval method to determine particular locations from the entirety of the selected content object 4816 that are relevant to a user query (e.g., a text query).In this example, the relevant parts (rather than the entirety of the information from the content object) are used by the one or more machine-learned models to generate a response to the user query (e.g., the text query).

[0116] Examples of the disclosure are directed to further user-facing aspects by which a user can manage content, organize content, create content, etc., via a notebook application configured to implement one or more machine-learned models with respect to source content selected by the user. For example, the Fig. 5A to 5B illustrate examples of actions that may be implemented for a project in which a note is generated via one or more machine-learned models based on source content selected by a user, in accordance with one or more example embodiments of the disclosure.

[0117] Fig. For example, Figure 5A illustrates a first user interface screen of a notebook application according to one or more example embodiments of the disclosure. For example, the first user interface screen 5100 in Fig. 5A, a first section 5110, a second section 5120, and a third section 5130. The first section 5110 corresponds to a notes sub-area (e.g., a scratchpad) that may include one or more notes 5112 that may be generated via various methods as described herein (e.g., automatically generated by the notebook application 3100, manually entered by a user, automatically generated by the notebook application 3100 in response to the selection of a user interface element corresponding to an action to be performed, etc.).

[0118] The second section 5120 corresponds to a source content portion (e.g., a source guide or context window) that may include one or more sources 5122 (e.g., content objects that comprise the source content 3400 that the one or more machine-learned models use to generate the information included in the one or more notes 5112). In some implementations, the notebook application 3100 may be configured to generate a note based on a selection of at least a portion of the information from a content object provided in the second section 5120, which is stored as a note in the first section 5110. For example, Fig. 5A selected text 5124 (e.g., highlighted text) selected by a user.

[0119] For example, the third section 5130 corresponds to a query portion, which may include one or more user interface elements for sending or providing a query to the notebook application 3100 regarding the source content. For example, the third section 5130 includes a plurality of user interface elements 5132 corresponding to suggested questions or actions related to the source content. For example, the notebook application 3100 may be configured to generate the suggested questions or actions based on information included in the source content and / or based on the information displayed in the second section 5120. The third section 5130 may further include a text input field 5134 with which a user can provide input (e.g., via a keyboard, via voice input, etc.) to query the notebook application 3100.The third section 5130 may further include a user interface element 5136 that indicates the number of content objects that comprise the source content. For example, in . Fig. 5A indicates that the notebook application 3100 used three sources to generate the one or more notes 5112.

[0120] In some implementations, the plurality of user interface elements 5132 may be configured to dynamically change based on actions with respect to the first user interface screen 5100. For example, the notebook application 3100 may be configured to dynamically change, modify, delete, or add user interface elements in the third section 5130 based on an action with respect to the source content (e.g., with respect to content objects provided for display in the second section 5120). Fig. 5A, the notebook application 3100 may be configured to dynamically change user interface elements in the third section 5130 based on (in response to) the selection of text from one or more sources 5122 (e.g., the selected text 5124). For example, as shown in Fig. 5A, the actions may include summarizing the selected text into a note, adding a quote to a note, requesting additional information regarding the selected text 5124, or suggesting related ideas. For example, the notebook application 3100 may be configured to generate a note summarizing the selected text in response to receiving a selection of the user interface element 5132a corresponding to the action of summarizing the selected text into a note. For example, the notebook application 3100 may be configured to add content to an existing note corresponding to the selected text in response to receiving a selection of the user interface element 5132b corresponding to the action of adding a quote to a note.

[0121] Fig. For example, Figure 5B illustrates a second user interface screen of a notebook application according to one or more example embodiments of the disclosure. For example, Fig. 5B, the second user interface screen 5200 includes a first section 5210, a second section 5220, and a third section 5230, each of which corresponds to the first section 5110, the second section 5120, and the third section 5130 of Fig. 5A.

[0122] As with regard to Fig. 5A, the notebook application 3100 may be configured to generate a note summarizing the selected text in response to receiving a selection of the user interface element 5132a corresponding to the action of summarizing the selected text into a note. Fig. 5B illustrates the generated note 5214 stored in the first section 5210, which includes one or more notes 5212. Further, in some implementations, after the generated note 5214 is stored in the first section 5210, the plurality of user interface elements 5132 may be Fig. 5A may be configured to dynamically switch back to a previous state to the plurality of user interface elements 5232 included in Fig. 5B is shown.

[0123] Examples of the disclosure are directed to further user-facing aspects by which a user can manage content, organize content, create content, etc., via a notebook application configured to implement one or more machine-learned models with respect to source content selected by the user. For example, the Fig. 6A to 6B illustrate examples of actions that may be implemented for a project in which a note is generated via one or more machine-learned models based on source content selected by a user, in accordance with one or more example embodiments of the disclosure.

[0124] Fig. For example, Figure 6A illustrates a portion of a first user interface screen of a notebook application according to one or more example embodiments of the disclosure. For example, in Fig. 6A, a first portion 6110 and a second portion 6120 of a user interface screen are shown. The first portion 6110 corresponds to a notes sub-area (e.g., a scratchpad) that may include a plurality of notes that may be created via various methods as described herein (e.g., automatically generated by the notebook application 3100, manually entered by a user, automatically generated by the notebook application 3100 in response to the selection of a user interface element corresponding to an action to be performed, etc.). For example, the first portion 6110 may indicate how a particular note is created (e.g., as a saved response, as a written note written by a user, as a document generated from other notes, etc.).

[0125] For example, the second section 6120 corresponds to a query portion, which may include one or more user interface elements for sending or providing a query to the notebook application 3100 regarding the source content or the plurality of notes. For example, the second section 6120 includes a plurality of user interface elements 6122 corresponding to suggested questions or actions related to the source content or the plurality of notes. For example, the notebook application 3100 may be configured to generate the suggested questions or actions based on information included in the source content and / or based on the information displayed in the first section 6110. The second section 6120 may further include a text input field with which a user can provide input (e.g., via a keyboard, via voice input, etc.).) to query the notebook application 3100, a user interface element indicating the number of content objects that comprise the source content.

[0126] In the example from Fig. 6A, the notebook application 3100 may be configured to dynamically change user interface elements in the second section 6120 based on the selection of one or more notes 6112 from the plurality of notes provided in the first section 6110 (in response thereto). For example, as in Fig. 6A, one or more notes may be selected via user input (e.g., via drag input, via selecting checkboxes, etc.), and in response to selecting the one or more notes, the actions may include actions to create content (e.g., creating a study guide, creating an outline, creating a spreadsheet, creating a presentation, etc.), suggesting related ideas, etc. based on the selected notes 6114. For example, the notebook application 3100 may be configured to generate a note corresponding to an outline of the content from the selected notes 6114 in response to receiving a selection of the user interface element 6122a corresponding to the action of creating an outline and saving the outline to a note.For example, in response to receiving a selection of a user interface element corresponding to an action of creating content related to the selected one or more notes and saving the content as a note, the notebook application 3100 may be configured to implement one or more machine-learned models to generate the note corresponding to the selected notes 6114.

[0127] Fig. For example, Figure 6B illustrates a portion of a second user interface screen of a notebook application according to one or more example embodiments of the disclosure. For example, in Fig. 6B the first section 6210 the first section 6110 from Fig. 6A.

[0128] As with regard to Fig. 6A, the notebook application 3100 may be configured to generate a note based on one or more selected notes by implementing one or more machine-learned models, where the selected notes may correspond to source content (e.g., source content selected by a user and used as input to generate the note). For example, the generated note may summarize or organize the notes described with respect to Fig. 6A. The notebook application 3100 may be configured to generate the generated note 6214 by implementing one or more machine-learned models based on the selected notes 6114, where the selected notes 6114 may correspond to source content (e.g., source content selected by a user and used as input to generate the note), and in response to receiving a selection of a user interface element (e.g., user interface element 6122a) corresponding to an action to combine the selected notes 6114 into the generated note 6214. Fig. 6B illustrates the generated note 6214 stored in the first section 6210, which includes one or more other notes 6212. Further, in some implementations, after the generated note 6214 is stored in the first section 6210, the plurality of user interface elements 6122 may be Fig. 6A be configured to dynamically return to a previous state.

[0129] In some implementations, the notebook application 3100 may be configured to enable a generated note 6214 to be exported to other applications via the selection of a user interface element to send the document to another application (e.g., a word processing application, a presentation application, a spreadsheet application, a social media application, etc.). In some implementations, the notebook application 3100 may be configured to enable a generated note 6214 and / or content objects (e.g., source content 3400) to be shared with other users via the selection of a user interface element to share the document and / or source content with another user.

[0130] According to examples of the disclosure, the notebook application 3100 may be configured to generate output (e.g., an outline, a report, a summary, etc.) via one or more machine-learned models based on source content provided to the scratchpad application (e.g., by the user). The notebook application 3100 may be configured to allow a user to create different projects to complete different tasks. Each project may be configured to function similarly to a folder through which a user can store various information about each project. In some implementations, a single scratchpad may correspond to or be dedicated to a particular project. In some implementations, the notebook application 3100 may be configured to receive the source content as defined by the user.The notebook application 3100 may be configured to add, delete, or modify projects according to input received from a user. Each project may be provided with a default name, a user-provided name, or a name generated by the notebook application 3100 (e.g., via one or more machine-learned models) based on the information stored in the project (e.g., based on the source content).

[0131] Examples of the disclosure are directed to further user-facing aspects by which a user can manage content, organize content, create content, etc., via a notebook application configured to implement one or more machine-learned models with respect to source content selected by the user. For example, Fig. 7 examples of notebooks or projects that can be represented in a specific way so that a user can easily understand the content contained in the notebook or project.

[0132] In some implementations, the notebook application 3100 may be configured to automatically generate (e.g., using one or more machine-learning models, one or more generative machine-learning models, semantic retrieval technologies, etc.) a graphical image (e.g., an emoji, an icon, etc.) or a graphical animation corresponding to or representing the source content in response to the notebook application 3100 being provided with source content. In some implementations, the graphical image or graphical animation may be overlaid on a folder provided as a user interface element that, when selected, causes the folder to open and display the contents of the folder to the user. Additionally or alternatively, in some implementations, the notebook application 3100 may be configured to automatically generate (e.g., using one or more machine-learning models, one or more generative machine-learning models, semantic retrieval technologies, etc.) a graphical image (e.g., an emoji, an icon, etc.) or a graphical animation corresponding to or representing the source content in response to the notebook application 3100 being provided with source content.The text description may be overlaid on the folder, provided as a user interface element that, when selected, causes the folder to open and display the folder's contents to the user.

[0133] With reference to Fig. 7, the notebook application 3100 may include a user-specific sub-area 7100 in which various projects are stored in specific folders. For example, a first folder 7110 (e.g., default folder) may be represented by a default image 7112 and have a generic name 7114 (e.g., "Default Notebook"). For example, a second folder 7120 may be represented by a graphical image 7122 and have a textual description 7124 (e.g., "Revenue") generated by machine learning and representing or corresponding to content included in the second folder 7120. For example, the notebook application 3100 may be configured to automatically (e.g., using one or more machine learning models, one or more generative machine learning models, semantic retrieval technologies, etc.) generate user-specific content in response to source content being provided to the notebook application 3100.) to generate the graphical image 7122, which may correspond to an emoji, an icon, etc., that corresponds to or represents the source content. In some implementations, the graphical image 7122 may be overlaid on the second folder 7120, which is provided as a user interface element that, when selected, causes the second folder 7120 to open and display the contents of the second folder 7120 to the user. Additionally or alternatively, in some implementations, the notebook application 3100 may be configured to automatically generate (e.g., using one or more machine-learning models, one or more generative machine-learning models, semantic retrieval technologies, etc.) the textual description 7124 (name) that corresponds to or represents the source content, in response to the notebook application 3100 being provided with the source content.The text description 7124 may be overlaid on the second folder 7120, which is provided as a user interface element that, when selected, causes the second folder 7120 to open and display the contents of the second folder 7120 to the user.

[0134] Fig. 8A illustrates a block diagram of an example computing system for organizing, managing, and creating content by implementing one or more machine-learned models with respect to user-selected source content, in accordance with one or more example embodiments of the disclosure. System 8100 includes a user computing device 8102, a server computing system 8130, and a training computing system 8150 communicatively coupled via a network 8180.

[0135] Fig. 8B illustrates a block diagram of an example computing device for organizing, managing, and creating content by implementing one or more machine-learned models with respect to user-selected source content, in accordance with one or more example embodiments of the disclosure.

[0136] Fig. 8C illustrates a block diagram of an example computing device for organizing, managing, and creating content by implementing one or more machine-learned models with respect to user-selected source content, in accordance with one or more example embodiments of the disclosure.

[0137] The user computing device 8102 (which may correspond to the computing device 100) may be any type of computing device, such as a personal computing device (e.g., laptop or desktop), a mobile computing device (e.g., smartphone or tablet), a gaming console or controller, a portable computing device, an embedded computing device, or any other type of computing device.

[0138] The user computing device 8102 includes one or more processors 8112 and a memory 8114. The one or more processors 8112 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or a plurality of processors operatively connected. The memory 8114 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 8114 may store data 8116 and instructions 8118 that are executed by the processor 8112 to cause the user computing device 8102 to perform operations.

[0139] In some implementations, the user computing device 8102 may store or include one or more machine-learned models 8120 (e.g., large language models, sequence processing models, generative machine-learned models, etc.). For example, the one or more machine-learned models 8120 may be or otherwise include various machine-learned models such as neural networks (e.g., deep neural networks) or other types of machine-learned models, including nonlinear models and / or linear models. Example neural networks may include feedforward neural networks, recurrent neural networks (RNNs), including long short-term memory (LSTM)-based recurrent neural networks, convolutional neural networks (CNNs), diffusion models, generative-adversarial networks, or other forms of neural networks.Examples of neural networks include deep neural networks. Some examples of machine-learned models may employ an attention mechanism, such as self-attention. For example, some machine-learned models may incorporate multi-head self-attention models (e.g., Transformer models). With reference to the . Fig. 1A to 7, exemplary machine-learned models were described in this paper.

[0140] In some implementations, the one or more machine-learned models 8120 may be received from the server computing system 8130 over the network 8180, stored in the memory 8114, and then used or otherwise implemented by the one or more processors 8112. In some implementations, the user computing device 8102 may implement multiple parallel instances of a single machine-learned model (e.g., to perform parallel tasks across multiple instances of the machine-learned model). In some implementations, the task is a generative task, and one or more machine-learned models may be implemented to output content (e.g., an answer to a question, a summary of various selected content objects, an outline of various selected notes, etc.) in terms of various inputs (e.g., a query, conditioning parameters, etc.).In particular, the machine-learned models disclosed in this document (e.g., including large language models, sequence processing models, generative machine-learned models, etc.) can be implemented to perform various tasks involving an input query.

[0141] According to examples of the disclosure, a computing system may implement one or more sequence processing models 3120, as described herein, to output values ​​for the one or more conditions in response to or based on the query. The one or more sequence processing models 3120 may include one or more machine-learned models configured to process and analyze sequential data and to process data that occurs in a particular order or sequence, including time series data, natural language text, or other data with a temporal or sequential structure.

[0142] According to examples of the disclosure, a computing system may implement one or more large language models 3130 to determine a plurality of variables based on the query. For example, a large language model may include Bidirectional Encoder Representations from Transformers (BERT). The large language model may, for example, be trained to understand and process natural language. The large language model may be configured to extract information from the input (query) to identify keywords, intents, and context within the input to determine a plurality of variables for generating content. The variables may include latent variables that represent an underlying structure of the language.

[0143] According to examples of the disclosure, a computing system may implement one or more generative machine-learned models 3140 to generate various content (e.g., to generate an outline, a summary, a response to a query, etc.) having values ​​for one or more conditions. The one or more generative machine-learned models 3140 may include a deep neural network or a generative adversarial network (GAN) to generate the content with one or more features having values ​​for one or more conditions associated with the features. For example, the one or more generative machine-learned models 3140 may include variational autoencoders, stable diffusion machine-learned models, visual transformers, neural radiance fields (NeRF), etc., to generate the content.

[0144] Additionally or alternatively, one or more machine-learned models 8140 may be included in, or otherwise stored and implemented by, the server computing system 8130, which communicates with the user computing device 8102 according to a client-server relationship. For example, the one or more machine-learned models 8140 may be implemented by the server computing system 8130 as part of a web service (e.g., a navigation service, a word processing service, an educational service, and the like). Thus, one or more machine-learned models 8120 may be stored and implemented on the user computing device 8102 and / or one or more machine-learned models 8140 may be stored and implemented on the server computing device 8130.

[0145] The user computing device 8102 may also include one or more user input components 8122 that receive user input. For example, the user input component 8122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or a touchpad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component may implement a virtual keyboard. Other example user input components include a microphone, a conventional keyboard, or other devices and methods that allow a user to provide user input.

[0146] The server computing system 8130 (which may correspond to the server computing system 300) includes one or more processors 8132 and memory 8134. The one or more processors 8132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or a plurality of processors operatively connected. The memory 8134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 8134 may store data 8136 and instructions 8138 that are executed by the processor 8132 to cause the server computing system 8130 to perform operations.

[0147] In some implementations, server computing system 8130 includes or is otherwise implemented by one or more server computing devices. In cases where server computing system 8130 includes a plurality of server computing devices, such server computing devices may operate according to sequential computing architectures, parallel computing architectures, or a combination thereof.

[0148] As described above, the server computing system 8130 may store or otherwise include one or more machine-learned models 8140. For example, the one or more machine-learned models 8140 may be or otherwise include different machine-learned models. Example machine-learned models include neural networks or other multi-layer nonlinear models. Example neural networks may include feedforward neural networks, recurrent neural networks (RNNs), including long short-term memory (LSTM)-based recurrent neural networks, convolutional neural networks (CNNs), diffusion models, generative adversarial networks, or other forms of neural networks. Examples of neural networks may include deep neural networks.Some examples of machine-learned models may employ an attention mechanism, such as self-attention. For example, some machine-learned models may incorporate multi-head self-attention models (e.g., Transformer models). With reference to the . Fig. 1A to 7, exemplary machine-learned models were described in this paper.

[0149] The user computing device 8102 and / or the server computing system 8130 may train the one or more machine-learned models 8120 and / or 8140 via interaction with the training computing system 8150 communicatively coupled via the network 8180. The training computing system 8150 may be separate from the server computing system 8130 or may be a portion of the server computing system 8130.

[0150] The training computing system 8150 includes one or more processors 8152 and a memory 8154. The one or more processors 8152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or a plurality of processors operatively connected. The memory 8154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 8154 may store data 8156 and instructions 8158 that are executable by the processor 8152 to cause the training computing system 8150 to perform operations. In some implementations, the training computing system 8150 includes or is otherwise implemented by one or more server computing devices.

[0151] The training computing system 8150 may include a model trainer 8160 that trains the one or more machine-learned models 8120 and / or 8140 stored on the user computing device 8102 and / or the server computing system 8130 using various training or learning techniques, such as backpropagation of errors. For example, a loss function may be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function). Various loss functions may be used, such as mean square error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques may be used to iteratively update the parameters over a number of training iterations.

[0152] In some implementations, performing error backpropagation may involve performing a truncated backpropagation over time. The model trainer 8160 may perform a variety of generalization techniques (e.g., weight decay, dropouts, etc.) to improve the generalization ability of the models being trained.

[0153] In particular, the model trainer 8160 may train the one or more machine-learned models 8120 and / or 8140 based on a training data set 8162. The training data 8162 may, for example, include various data sets that may be remote or stored on the training computing system 8150. In some implementations, for example, a sample data set used for training includes a large corpus of language training data that provides one or more large language models with the ability to perform multiple language tasks. For example, the one or more large language models may be trained to perform summarization tasks, conversation tasks, simplification tasks, opposing viewpoint tasks, etc. In particular, the one or more large language models may be trained to process a variety of outputs to produce a speech output.However, other datasets (e.g., images obtained from external websites) may also be used. In some implementations, the dataset may be limited to a specific genre or topic, specific types of content (including images, videos, and text), specific types of content (including outlines, reports, presentations, spreadsheets, etc.), etc. In some implementations, the dataset may contain various objects.

[0154] In some implementations, the training examples may be provided by the user computing device 8102 if the user has provided consent. Thus, in such implementations, the one or more machine-learned models 8120 provided to the user computing device 8102 may be trained by the training computing system 8150 with user-specific data received from the user computing device 8102. In some cases, this process may be referred to as personalizing the model.

[0155] The model trainer 8160 includes computer logic employed to provide the desired functionality. The model trainer 8160 may be implemented in hardware, firmware, and / or software controlling a general-purpose processor. For example, in some implementations, the model trainer 8160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, the model trainer 8160 includes one or more sets of computer-executable instructions stored in a physical computer-readable storage medium such as RAM, hard disk, or optical or magnetic media.

[0156] The network 8180 may be any type of communications network, such as a local area network (e.g., intranet), a wide area network (e.g., Internet), or a combination thereof, and may include any number of wired or wireless connections. In general, communication over the network 8180 may occur over any type of wired and / or wireless connection, using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).

[0157] The machine-learned models described in this description can be used in a variety of tasks, applications, and / or use cases.

[0158] In some implementations, the input to the machine-learned model(s) of the disclosure may be text data or natural language data. The machine-learned model(s) may process the text data or natural language data to produce an output. As one example, the machine-learned model(s) may process the natural language data to produce a speech coding output. As another example, the machine-learned model(s) may process the text data or natural language data to produce a latent text embedding output. As another example, the machine-learned model(s) may process the text data or natural language data to produce a translation output. As another example, the machine-learned model(s) may process the text data or natural language data to produce a classification output.As another example, the machine-learned model(s) may process the text data or natural language data to produce a text segmentation output. As another example, the machine-learned model(s) may process the text data or natural language data to produce a semantic intent output. As another example, the machine-learned model(s) may process the text data or natural language data to produce an upscaled text or natural language output (e.g., text data or natural language data that is of higher quality than the input text or natural language, etc.). As another example, the machine-learned model(s) may process the text data or natural language data to produce a prediction output.

[0159] In some implementations, the input to the machine-learned model(s) of the disclosure may be speech data. The machine-learned model(s) may process the speech data to produce an output. As one example, the machine-learned model(s) may process the speech data to produce a speech recognition output. As another example, the machine-learned model(s) may process the speech data to produce a speech translation output. As another example, the machine-learned model(s) may process the speech data to produce a latent embedding output. As another example, the machine-learned model(s) may process the speech data to produce an encoded speech output (e.g., an encoded and / or compressed representation of the speech data, etc.).As another example, the machine-learned model(s) may process the speech data to produce an upscaled speech output (e.g., speech data that is of higher quality than the input speech data, etc.). As another example, the machine-learned model(s) may process the speech data to produce a text representation output (e.g., a text representation of the input speech data, etc.). As another example, the machine-learned model(s) may process the speech data to produce a prediction output.

[0160] In some implementations, the input to the machine-learned model(s) of the disclosure may be sensor data. The machine-learned model(s) may process the sensor data to produce an output. As one example, the machine-learned model(s) may process the sensor data to produce a detection output. As another example, the machine-learned model(s) may process the sensor data to produce a prediction output. As another example, the machine-learned model(s) may process the sensor data to produce a classification output. As another example, the machine-learned model(s) may process the sensor data to produce a segmentation output. As another example, the machine-learned model(s) may process the sensor data to produce a visualization output.As another example, the machine-learned model(s) may process the sensor data to generate a diagnostic output. As another example, the machine-learned model(s) may process the sensor data to generate a detection output.

[0161] Fig. 8A illustrates an example computing system that may be used to implement aspects of the disclosure. Other computing systems may also be used. For example, in some implementations, the user computing device 8102 may include the model trainer 8160 and the training data 8162. In such implementations, the one or more machine-learned models 8120 may be both trained and used locally at the user computing device 8102. In some such implementations, the user computing device 8102 may implement the model trainer 8160 to personalize the one or more machine-learned models 8120 based on user-specific data.

[0162] Fig. 8B illustrates a block diagram of an example computing device for organizing, managing, and creating content by implementing one or more machine-learned models with respect to user-selected source content, in accordance with one or more example embodiments of the disclosure. Computing device 8200 may be a user computing device or a server computing device.

[0163] Computing device 8200 includes a number of applications (e.g., applications 1 through N). Each application includes its own machine learning library and machine-learned model(s). For example, each application may include a machine-learned model. Example applications include a notebook application, a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, a social media application, a mapping application, a navigation application, etc.

[0164] As in Fig. As illustrated in Figure 8B, each application may communicate with a number of other components of the computing device, such as one or more sensors, a context manager, a device state component, or additional components. In some implementations, each application may communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.

[0165] Fig. 8C illustrates a block diagram of an example computing device for organizing, managing, and creating content by implementing one or more machine-learned models with respect to user-selected source content, in accordance with one or more example embodiments of the disclosure. Computing device 80 may be a user computing device or a server computing device.

[0166] Computing device 8300 includes a number of applications (e.g., applications 1 through N). Each application communicates with a central intelligence layer. Example applications include a notebook application, as described herein, a notepad application, a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, a mapping application, a social media application, a navigation application, a social media application, etc. In some implementations, each application can communicate with the central intelligence layer (and the model(s) stored therein) using an API (e.g., a common API for all applications).

[0167] The central intelligence layer includes a number of machine-learned models. For example, as in Fig. 8C illustrates that a respective machine-learned model may be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications may share a single machine-learned model. For example, in some implementations, the central intelligence layer may provide a single model for all applications. In some implementations, the central intelligence layer is included in or otherwise implemented by an operating system of the computing device 8300.

[0168] The central intelligence layer may communicate with a central device data layer. The central device data layer may be a central repository of data for the computing device 8300. As in Fig.As illustrated in Figure 8C, the central device data layer may communicate with a number of other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer may communicate with each device component using an API (e.g., a private API).

[0169] Where generic terms such as "module," "unit," and the like are used in this document, these terms may refer to, but are not limited to, a software or hardware component or device, such as a Field Programmable Gate Array (FPGA) or Application Specific Integrated Circuit (ASIC), that performs specific tasks. A module or unit may be configured to reside on an addressable storage medium and configured to execute on one or more processors. Thus, a module or unit may include, for example, components such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables.The functionality provided in the components and modules / units can be combined into fewer components and modules / units or further divided into additional components and modules.

[0170] Aspects of the exemplary embodiments described above may be recorded on non-transitory computer-readable media, including program instructions for implementing various operations implemented by a computer. The media may also include, alone or in combination with the program instructions, data files, data structures, and the like. Examples of non-transitory computer-readable media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CD-ROM disks, Blu-ray disks, and DVDs; magnetic-optical media such as optical disks; and other hardware devices specifically configured to store and execute program instructions, such as semiconductor memory, read-only memory (ROM), random access memory (RAM), flash memory, USB storage, and the like.Examples of program instructions include both machine code, such as that generated by a compiler, and files containing higher-level code that can be executed by the computer using an interpreter. The program instructions can be executed by one or more processors. The described hardware devices can be configured to function as one or more software modules to perform the operations of the embodiments described above, or vice versa. Furthermore, a non-transitory computer-readable storage medium can be distributed among computer systems connected via a network, and computer-readable code or program instructions can be stored and executed in a decentralized manner.In addition, the non-transitory computer-readable storage media can also be implemented in at least one application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).

[0171] Each block of the flowchart illustrations may represent a unit, module, segment, or section of code that includes one or more executable instructions for implementing the defined logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order. For example, two consecutive blocks shown may actually execute substantially simultaneously, or the blocks may sometimes execute in the reverse order depending on the particular function.

[0172] While the disclosure has been described with respect to various exemplary embodiments, each example is provided for the purpose of illustration, not limitation. Those skilled in the art, after receiving the foregoing, may readily make modifications, variations, and equivalents to such embodiments. Accordingly, the disclosure does not preclude the inclusion within the disclosed subject matter of such modifications, variations, or additions as would be readily apparent to those skilled in the art. For example, features illustrated or described as part of one embodiment may be used in another embodiment to yield yet another embodiment. Therefore, the disclosure is intended to cover such modifications, variations, and equivalents.

Claims

[1] Computing device for generating content, comprising: one or more memories configured to store instructions; and one or more processors configured to execute the instructions to perform operations, the operations comprising: Providing, in response to a selection of a plurality of content objects, a user interface including a first portion and a second portion, wherein the first portion includes a summary description generated via one or more machine-learned models based on the plurality of content objects, and the second portion includes a plurality of user interface elements configured to perform an operation with respect to at least one of the summary description or the plurality of content objects. [2] The computing device of claim 1, wherein the operations further comprise: Receiving a selection of the plurality of content objects; and Implementing the one or more machine-learned models to generate the summary description based on the plurality of content objects. [3] Computing device according to claim 1, wherein the plurality of user interface elements includes a first user interface element comprising a suggested query, and the operations further comprising: Receiving a selection of the first user interface element; and Implementing the one or more machine-learned models to generate a response to the suggested query based on at least one of the summary description or the plurality of content objects. [4] The computing device of claim 3, wherein the operations further comprise providing a third portion for the user interface to provide a dialog for display including the suggested query and the response. [5] The computing device of claim 4, wherein the third portion includes a quote user interface element indicating a number of content objects from the plurality of content objects referenced by the one or more machine-learned models to generate the response. [6] The computing device of claim 5, wherein the operations further comprise: in response to receiving a selection of the quote user interface element, providing, for display in a fourth portion of the user interface, content from one or more content objects of the plurality of content objects used to generate the response. [7] The computing device of claim 4, wherein the third portion includes a note creation user interface element, and wherein the operations further comprise: in response to receiving a selection of the note creation user interface element, creating a note that includes content from the suggested query and the answer, and Save the note. [8] Computing device according to claim 1, wherein the plurality of user interface elements includes a first user interface element configured to generate a note, and the operations further comprising: Providing a third portion for the user interface that includes content from one or more content objects from the plurality of content objects; Receiving a selection of a portion of the content from the one or more content objects from the plurality of content objects; Receiving a selection of the first user interface element; and in response to receiving the selection of the section of content and the first user interface element, generating, via the one or more machine learned models based on the section of content, the note that includes a summary of the section of content. [9] Computing device according to claim 1, wherein the plurality of user interface elements includes a first user interface element configured to add content to an existing note, and the operations further comprising: Providing a third portion for the user interface that includes content from one or more content objects from the plurality of content objects; Receiving a selection of a portion of the content from the one or more content objects from the plurality of content objects; Receiving a selection of the first user interface element; and in response to receiving the selection of the section of content and the first user interface element, adding the section of content to the existing note. [10] Computing device according to claim 1, wherein the first section further includes at least one key topic user interface element comprising at least one key topic relating to the summary description of the plurality of content objects, and the operations further comprising receiving a selection of the at least one key theme user interface element; and Implementing the one or more machine-learned models to generate an output relating to the at least one key topic based on at least one of the summary description or the plurality of content objects. [11] The computing device of claim 10, wherein the operations further comprise providing a third portion for the user interface to provide a dialog for display including the key topic and the output related to the at least one key topic. [12] The computing device of claim 1, wherein the operations further comprise: Providing a third portion for the user interface to provide for display of at least one note generated via the one or more machine learned models based on the plurality of content objects. [13] The computing device of claim 12, wherein the third portion includes a quote user interface element indicating a number of content objects from the plurality of content objects referenced by the one or more machine-learned models to generate the at least one note. [14] The computing device of claim 13, wherein the operations further comprise: in response to receiving a selection of the quote user interface element, providing, for display in a fourth portion of the user interface, content from one or more content objects of the plurality of content objects used to create the note. [15] The computing device of claim 14, wherein the operations further comprise: in response to receiving the selection of the quote user interface element, providing, for display in the fourth portion of the user interface, contextual content about the content from the one or more content objects from the plurality of content objects used to generate the note. [16] The computing device of claim 12, wherein the plurality of user interface elements includes a first user interface element configured to generate new content based on one or more notes, and wherein the operations further comprise: Providing a third portion for the user interface to provide a plurality of notes for display generated via the one or more machine learned models based on the plurality of content objects; Receiving a selection of the plurality of notes; Receiving a selection of the first user interface element; and in response to receiving the selection of the plurality of notes and the first user interface element, generating the new content based on the plurality of notes. [17] The computing device of claim 1, wherein the operations further comprise: Generating, via the one or more machine-learned models, a graphical image representing the plurality of content objects; and Providing a folder containing the graphic image, the folder storing the plurality of content objects and a project file containing the summary description. [18] Computing device according to claim 1, wherein the second section includes a text input field for receiving a query from a user, and the operations further comprising: Implementing the one or more machine-learned models to generate a response to the query based on at least one of the summary description or the plurality of content objects. [19] Computing device for generating content, comprising: one or more memories configured to store instructions; and one or more processors configured to execute the instructions to perform operations, the operations comprising: Receiving input to create a notebook; Receiving a selection of a plurality of content objects to be added to the notebook; in response to receiving the selection of the plurality of content objects, implementing one or more machine-learned models to generate a summary description based on the plurality of content objects and at least one of a key topic user interface element indicating a topic of the plurality of content objects or a selectable user interface element indicating a query concerning the plurality of content objects; and Providing a user interface including a first section and a second section, wherein the first section includes the summary description and the second section includes the selectable user interface element. [20] Computing device according to claim 19, wherein the second section includes a text input field for receiving a query from a user, and the operations further comprising: Implementing the one or more machine-learned models to generate a response to the query based on at least one of the summary description or the plurality of content objects. [21] Computing device according to claim 19, wherein the selectable user interface element indicating the query concerning the plurality of content objects comprises a suggested query, and the operations further comprising: Receiving a selection of the selectable user interface element; and Implementing the one or more machine-learned models to generate a response to the suggested query based on at least one of the summary description or the plurality of content objects. [22] The computing device of claim 21, wherein the operations further comprise providing a third portion for the user interface to provide a dialog for display including the suggested query and the response. [23] The computing device of claim 22, wherein the third portion includes a quote user interface element indicating a number of content objects from the plurality of content objects referenced by the one or more machine-learned models to generate the response. [24] The computing device of claim 23, wherein the operations further comprise: in response to receiving a selection of the quote user interface element, providing, for display in a fourth portion of the user interface, content from one or more content objects of the plurality of content objects used to generate the response. [25] The computing device of claim 22, wherein the third portion includes a note creation user interface element, and wherein the operations further comprise: in response to receiving a selection of the note creation user interface element, creating a note that includes content from the suggested query and the answer, and Save the note. [26] Computing device according to claim 19, wherein the user interface includes a note creation user interface element configured to create a note, and the operations further comprising: Providing a third section on the user interface that includes content from one or more content objects from the plurality of content objects; Receiving a selection of a portion of the content from the one or more content objects from the plurality of content objects; Receiving a selection of the note creation user interface element; and in response to receiving the selection of the section of content and the note generation user interface element, generate, via the one or more machine learned models based on the section of content, the note that includes a summary of the section of content. [27] Computing device according to claim 19, wherein the user interface includes an editing user interface element configured to add content to an existing note, and the operations further comprising: Providing a third section on the user interface that includes content from one or more content objects from the plurality of content objects; Receiving a selection of a portion of the content from the one or more content objects from the plurality of content objects; Receiving a selection of the editing user interface element; and in response to receiving the selection of the section of content and the editing user interface element, adding the section of content to the existing note. [28] Computing device according to claim 19, wherein the operations further comprising receiving a selection of the key theme user interface element; and Implementing the one or more machine-learned models to generate an output relating to the key topic based on at least one of the summary description or the plurality of content objects. [29] The computing device of claim 28, wherein the operations further comprise providing a third portion on the user interface to provide a dialog for display including the topic and the output related to the topic. [30] The computing device of claim 19, wherein the operations further comprise: Providing a third portion on the user interface to provide for display at least one note generated via the one or more machine learned models based on the plurality of content objects. [31] The computing device of claim 30, wherein the third portion includes a quote user interface element indicating a number of content objects from the plurality of content objects referenced by the one or more machine-learned models to generate the at least one note. [32] The computing device of claim 31, wherein the operations further comprise: in response to receiving a selection of the quote user interface element, providing, for display in a fourth portion of the user interface, content from one or more content objects of the plurality of content objects used to create the note. [33] The computing device of claim 32, wherein the operations further comprise: in response to receiving the selection of the quote user interface element, providing, for display in the fourth portion of the user interface, contextual content about the content from the one or more content objects from the plurality of content objects used to generate the note. [34] The computing device of claim 30, wherein the user interface includes a content generation user interface element configured to generate new content based on one or more notes, and wherein the operations further comprise: Providing a third section on the user interface to provide a plurality of notes for display generated via the one or more machine-learned models based on the plurality of content objects; Receiving a selection of the plurality of notes; Receiving a selection of the content generation user interface element; and in response to receiving the selection of the plurality of notes and the content generation user interface element, generating the new content based on the plurality of notes. [35] The computing device of claim 19, wherein the operations further comprise: Generating, via the one or more machine-learned models, a graphical image representing the plurality of content objects; and Providing a folder containing the graphic image, the folder storing the plurality of content objects and a project file containing the summary description.