Machine learning-driven content generation via predictive content generation space

The predictive content generation space integrates machine-learned models to facilitate efficient content generation and discovery by allowing users to interact with multiple tools within a unified interface, addressing the inefficiencies of separate service navigation.

JP2025532817APending Publication Date: 2025-10-03GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025517449
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-09-23
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Current machine-learned models are underutilized for brainstorming, content discovery, and creative exploration due to users needing to navigate separate services for specific tasks, leading to inefficient use of computational resources.

Method used

A predictive content generation space utilizing machine-learned models allows users to interact with various tools within a continuous interface, enabling efficient content generation and discovery by selecting content elements and processing them with models trained for specific tasks.

Benefits of technology

This approach reduces computational resource consumption and enhances user efficiency, creativity, and productivity by integrating multiple machine learning tasks within a unified space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025532817000001_ABST
    Figure 2025532817000001_ABST
Patent Text Reader

Abstract

Systems and methods for generating content are provided. The method includes obtaining data indicative of user selections of content elements to be depicted within the predictive content generation space using tools of a predictive content generation space. Each tool is associated with a machine learning task. The tools are operable to select at least a portion of each of one or more content elements to be depicted within the predictive content generation space. The method includes processing data describing at least a portion of the content elements with a machine-learned model to obtain predictive content. The machine-learned model is trained to perform the machine learning task associated with the tool. The method includes generating one or more predictive content elements within the predictive content generation space. The one or more predictive content elements describe the predictive content.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates generally to content generation and prediction, and more particularly to utilizing a predictive content generation space to generate content using machine-learned models. [Background technology]

[0002] Advances in machine learning are leading to the creation of increasingly sophisticated machine-learned models. For example, large-scale language models (LLMs) are trained with significantly larger amounts of data than those used to train traditional language models. In doing so, LLMs can be trained to perform multiple natural language processing tasks. In another example, image processing models can be trained to perform semantic image analysis of images. In other words, in addition to identifying objects depicted in an image, these image processing models can gain a semantic understanding of the scene itself.

[0003] Currently, many of these models are used to help users perform various tasks. For example, LLMs may be used to answer questions posed by users of a search service. In other examples, image processing models may be used to perform reverse image searches or to provide image suggestions to users. However, in current implementations, users only use these models after identifying a problem, preventing them from being effectively used for brainstorming, content discovery, content generation, creative exploration, etc. Summary of the Invention

[0004] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the description that follows, or may be learned from the description, or may be learned by practice of the embodiments.

[0005] One exemplary aspect of the present disclosure is directed to a computer-implemented method for generating content within a predictive content generation space via user-specified machine learning tasks. The method includes obtaining, by a computing system including one or more computing devices, data indicative of a user's selection of at least a portion of content elements to be depicted within the predictive content generation space using a first tool among a plurality of tools of the predictive content generation space. The plurality of tools are associated with a plurality of machine learning tasks, and each of the plurality of tools is operable to select at least a portion of each of one or more content elements to be depicted within the predictive content generation space. The method includes processing, by the computing system, data describing at least a portion of the content elements with machine-learned models to obtain predictive content, the machine-learned models being trained to perform a first machine learning task associated with the first tool. The method includes generating, by the computing system, one or more predictive content elements within the predictive content generation space, the one or more predictive content elements describing the predictive content.

[0006] Another exemplary aspect of the present disclosure is directed to a computing system for generating content within a predictive content generation space through user-specified machine learning tasks. The computing system includes one or more processors. The computing system includes one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations include obtaining data indicative of a user's selection of at least a portion of content elements to be depicted within the predictive content generation space using a first tool of a plurality of tools in the predictive content generation space. The plurality of tools are associated with multiple machine learning tasks, respectively. Each of the plurality of tools is operable to select at least a portion of each of the one or more content elements to be depicted within the predictive content generation space. The operations include processing data describing at least a portion of the content elements with machine-learned models to obtain predictive content, the machine-learned models each being trained to perform a first machine learning task associated with the first tool. The operations include obtaining one or more predictive content elements within the predictive content generation space, the one or more predictive content elements describing the predictive content.

[0007] Another exemplary aspect of the present disclosure is directed to one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations. The operations include using a first tool of a plurality of tools of a predictive content generation space to obtain data indicative of a user's selection of at least a portion of content elements to be depicted in the predictive content generation space. The plurality of tools are associated with a plurality of machine learning tasks, respectively. Each of the plurality of tools is operable to select at least a portion of each of the one or more content elements to be depicted in the predictive content generation space. The operations include processing data describing at least a portion of the content elements with machine-learned models to obtain predictive content, the machine-learned models each being trained to perform a first machine learning task associated with the first tool. The operations include selecting one or more predictive content elements in the predictive content generation space, the one or more predictive content elements describing the predictive content.

[0008] Other aspects of the present disclosure are directed to various systems, apparatus, non-transitory computer-readable media, user interfaces, and electronic devices.

[0009] These and other features, aspects, and advantages of various embodiments of the present disclosure will become better understood with reference to the following description and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the detailed description, serve to explain associated principles.

[0010] Detailed descriptions of embodiments directed to those skilled in the art are set forth herein with reference to the accompanying drawings. [Brief explanation of the drawings]

[0011] [Figure 1A]1 illustrates a block diagram of an exemplary computing system for generating content within a predictive content generation space, in accordance with some implementations of the present disclosure. [Figure 1B] 1 illustrates a block diagram of an exemplary computing device for generating content within a predictive content generation space, in accordance with some implementations of the present disclosure. [Figure 1C] 1 illustrates a block diagram of an exemplary computing device for training machine-learned model(s) to generate predictive content, according to some implementations of the present disclosure. [Figure 2A] 1 illustrates an example layout of an interface for a predictive content generation space at a first time, according to some implementations of the present disclosure. [Figure 2B] 1 illustrates an example interface for machine-learned generation of predictive content elements within a predictive content generation space, according to some implementations of the present disclosure. [Figure 2C] 10 illustrates an example layout of an interface for a predictive content generation space at a second time, according to some implementations of the present disclosure. [Figure 2D] 10 illustrates content elements and predicted content elements within an interface of a predictive content generation space at a second time, according to some implementations of the present disclosure. [Figure 3A] 10 illustrates an example layout of an interface in which alternative user selection inputs are provided at a first time, according to some other embodiments of the present disclosure. [Figure 3B] 10 illustrates predicted content elements corresponding to a selection of an entire content element within an interface at a second time, according to some implementations of the present disclosure. [Figure 3C] 10 illustrates predicted content elements corresponding to a selection of a predicted content element in an interface at a third time, according to some implementations of the present disclosure. [Figure 4A]10 illustrates an example layout of an interface in which an audio brush is selected along with an audio utterance provided by a user at a first time, according to some other implementations of the present disclosure. [Figure 4B] 10 illustrates predicted content elements corresponding to a selection of a content element in an interface via an audio brush tool at a second time, according to some implementations of the present disclosure. [Figure 5] 1 illustrates a data structure that associates tools with machine learning tasks, according to some embodiments of the present disclosure. [Figure 6] 1 illustrates a block diagram of an example of a machine-learned model according to some embodiments of the present disclosure. [Figure 7] FIG. 1 illustrates a block diagram of an example of a machine-learned model according to some other embodiments of the present disclosure. [Figure 8] FIG. 1 illustrates a block diagram of an example of a machine-learned model according to some other embodiments of the present disclosure. [Figure 9] 1 illustrates a flowchart diagram of an exemplary method for generating content within a predictive content generation space, according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0012] Reference numbers repeated among the drawings are intended to identify like features in the various embodiments.

[0013] overview Generally, the present disclosure is directed to content generation and prediction. More specifically, the present disclosure relates to utilizing a predictive content generation space to generate content using machine-learned models. For example, a computing system (e.g., a user device, a server hosting a predictive content generation space service, etc.) can obtain data indicating that a user has selected some (or all) of the content elements to be depicted in the predictive content generation space using one of several tools. The predictive content generation space can be a two-dimensional or three-dimensional space (e.g., an augmented reality (AR) / virtual reality (VR) space, etc.) in which the content elements are depicted. The content elements can be interface elements including image(s), video data, text content, Uniform Resource Locators (URLs), audio data, etc. A user can interact with the content elements through various tools. Each of these tools can correspond to a different machine learning task. Thus, by selecting content elements with a particular brush (e.g., by drawing a line through the content elements with the brush), a user can indicate the machine learning task they want to perform.

[0014] Continuing with the previous example, a user may select a content element including an image of a cat with a content analysis brush corresponding to a content analysis task. A computing system may process data describing the content element (e.g., metadata associated with the image of the cat) using a machine-learned model trained to perform the content analysis task to obtain predictive content. For example, the predictive content may be textual content that describes information about the breed of the cat depicted in the image, provides clarifying prompts to the user corresponding to the image of the cat, etc.

[0015] The computing system can generate predictive content element(s) that describe the predictive content. For example, if the predictive content includes images similar to the input image and text content related to the breed of cat depicted in the input image, the computing system may generate a predictive content element that includes both the similar images and the text content in the predictive content element. In this manner, the predictive content generation space can be utilized in conjunction with advanced machine learning models to improve user efficiency, creativity, and productivity.

[0016] Embodiments of the present disclosure provide several technical effects and advantages. As one example of a technical effect and advantage, a user utilizing traditional model implementations for content generation (e.g., search engines, reverse image search services, language processing services, etc.) must navigate among various separate services configured to perform only narrow tasks. For example, a user may utilize a reverse image search service to find output images that are similar to an input image. If the user desires to obtain additional information about an entity depicted in the output images, the user is forced to store those output images locally while attempting to navigate to a different service for semantic image analysis, unnecessarily utilizing a significant amount of computational resources (e.g., power, memory, storage, bandwidth, compute cycles, etc.). However, embodiments of the present disclosure facilitate efficient content discovery and generation by utilizing advanced machine-learned models in combination with a continuous, predictive content generation space that provides users with a variety of tools. By providing a more efficient content generation space, embodiments of the present disclosure can significantly reduce the amount of computational resources consumed by users.

[0017] Furthermore, embodiments of the present disclosure may in fact provide alternative graphical shortcuts within the graphical user interface that allow a user to directly access and configure tools for a machine-learned model and specify inputs to the model.

[0018] Referring now to the drawings, exemplary embodiments of the present disclosure will be described in more detail.

[0019] Exemplary Devices and Systems 1A illustrates a block diagram of an exemplary computing system 100 for generating content within a predictive content generation space, in accordance with some implementations of the present disclosure. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150, communicatively coupled via a network 180.

[0020] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0021] The user computing device 102 includes one or more processors 112 and memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple operatively connected processors. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.

[0022] In some implementations, the user computing device 102 can store or include one or more machine-learned models 120. For example, the machine-learned models 120 can be or otherwise include various machine-learned models, such as neural networks (e.g., deep neural networks), or other types of machine-learned models, including nonlinear and / or linear models. The neural networks can include feedforward neural networks, recurrent neural networks (e.g., long-short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. Some exemplary machine-learned models can utilize attention mechanisms, such as self-attention. For example, some exemplary machine-learned models can include multi-head self-attention models (e.g., Transformer models). Exemplary machine-learned models 120 are described with reference to FIGS. 5-7.

[0023] In some implementations, one or more entire models 120 may be received from server computing system 130 over network 180, stored in user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, user computing device 102 may implement multiple parallel instances of a single machine-learned model 120 (e.g., to perform parallel machine learning tasks across multiple instances of machine-learned model(s) 120).

[0024] More specifically, machine-learned model(s) 120 can be one or more models trained to perform various machine learning tasks. For example, machine-learned model(s) 120 may be or otherwise include large-scale language models (LLMs) trained to perform various natural language processing tasks. In other examples, machine-learned model(s) 120 may include semantic image processing models trained to perform semantic image analysis tasks (e.g., depicted entity recognition, scene determination, etc.). Additionally or alternatively, in some implementations, machine-learned model(s) 120 may be or otherwise include a machine-learned model pipeline, ensemble, etc., including multiple machine-learned models configured to process inputs in a particular order (e.g., an order corresponding to the machine learning task).

[0025] For example, a content development task (e.g., a task of finding content similar to an input) may specify that a machine-learned semantic image processing model processes an input image to obtain a semantic description of the input image, and then that a machine-learned content retrieval model processes the semantic description to obtain predicted content that is similar to the input. Thus, it should be broadly understood that machine-learned model(s) 120 may be or otherwise include any type of grouping or collection of machine-learned model(s).

[0026] Additionally or alternatively, one or more machine-learned models 140 may be included in or otherwise stored and implemented by a server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the machine-learned models 140 may be implemented by the server computing system 130 as part of a web service (e.g., a predictive content generation spatial service). Thus, one or more models 120 may be stored and implemented at the user computing device 102 and / or one or more models 140 may be stored and implemented at the server computing system 130.

[0027] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). The touch-sensitive component may function to implement a virtual keyboard. Other exemplary user input components include a microphone, a conventional keyboard, or other means by which a user can provide user input.

[0028] The server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple operatively connected processors. The memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 can store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.

[0029] In some implementations, server computing system 130 includes or is otherwise implemented by one or more server computing devices. If server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or any combination thereof.

[0030] As described above, the server computing system 130 may store or otherwise include one or more machine-learned models 140. For example, the models 140 may be or otherwise include various machine-learned models. Exemplary machine-learned models include neural networks or other multi-layer nonlinear models. Exemplary neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some exemplary machine-learned models may utilize attention mechanisms such as self-attention. For example, some exemplary machine-learned models may include multi-head self-attention models (e.g., Transformer models). Exemplary models 140 are described with reference to FIGS. 5-7.

[0031] The user computing device 102 and / or the server computing system 130 can train the models 120 and / or 140 by interacting with a training computing system 150 that is communicatively coupled via a network 180. The training computing system 150 can be separate from the server computing system 130 or can be part of the server computing system 130.

[0032] Training computing system 150 includes one or more processors 152 and memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple operatively connected processors. Memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. Memory 154 can store data 156 and instructions 158 that are executed by processor 152 to cause training computing system 150 to perform operations. In some implementations, training computing system 150 includes or is otherwise implemented by one or more server computing devices.

[0033] The training computing system 150 may include a model trainer 160 that trains the machine-learned models 120 and / or 140 stored on the user computing device 102 and / or the server computing system 130 using various training or learning techniques, such as backpropagation. For example, a loss function may be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent may be used to iteratively update the parameters over several training iterations.

[0034] In some implementations, performing backpropagation may include performing truncated backpropagation through time. The model trainer 160 may perform several generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.

[0035] In particular, model trainer 160 can train machine-learned models 120 and / or 140 based on a set of training data 162. Training data 162 can include, for example, sufficient data to train an advanced model such as an LLM or a semantic image processing model (e.g., language data, image data, etc.).

[0036] In some implementations, if the user provides consent, the training examples may be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 may be trained by the training computing system 150 with user-specific data received from the user computing device 102. In some examples, this process may be referred to as personalizing the model.

[0037] Model trainer 160 includes computer logic utilized to provide desired functionality. Model trainer 160 can be implemented in hardware, firmware, and / or software controlling a general-purpose processor. For example, in some embodiments, model trainer 160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other embodiments, model trainer 160 includes one or more sets of computer-executable instructions stored on a tangible computer-readable storage medium, such as RAM, a hard disk, or an optical or magnetic medium.

[0038] Network 180 can be any type of communications network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or any combination thereof, and can include any number of wired or wireless links. Generally, communications over network 180 can be transmitted over any type of wired and / or wireless connection using a wide variety of communications protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or security schemes (e.g., VPN, Secure HTTP, SSL).

[0039] The machine-learned models described herein may be used in a variety of tasks, applications, and / or use cases.

[0040] In some implementations, the input to the machine-learned model(s) of the present disclosure can be image data. The machine-learned model(s) can process the image data to generate an output. As an example, the machine-learned model(s) can process the image data to generate an image recognition output (e.g., a recognition of the image data, a latent embedding of the image data, an encoded representation of the image data, a hash of the image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an image segmentation output. As another example, the machine-learned model(s) can process the image data to generate an image classification output. As another example, the machine-learned model(s) can process the image data to generate an image data modification output (e.g., a modification of the image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an encoded image data output (e.g., an encoded and / or compressed representation of the image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an upscaled image data output. As another example, the machine-learned model(s) can process the image data to generate a prediction output.

[0041] In some implementations, input to the machine-learned model(s) of the present disclosure can be text or natural language data. The machine-learned model(s) can process the text or natural language data to generate an output. As an example, the machine-learned model(s) can process the natural language data to generate a language-encoded output. As another example, the machine-learned model(s) can process the text or natural language data to generate a latent text embedding output. As another example, the machine-learned model(s) can process the text or natural language data to generate a translation output. As another example, the machine-learned model(s) can process the text or natural language data to generate a classification output. As another example, the machine-learned model(s) can process the text or natural language data to generate a text segmentation output. As another example, the machine-learned model(s) can process the text or natural language data to generate a semantic intent output. As another example, the machine-learned model(s) may process text or natural language data to generate upscaled text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language). As another example, the machine-learned model(s) may process text or natural language data to generate predicted outputs.

[0042] In some implementations, input to the machine-learned model(s) of the present disclosure can be speech data. The machine-learned model(s) can process the speech data to generate an output. As an example, the machine-learned model(s) can process the speech data to generate a speech recognition output. As another example, the machine-learned model(s) can process the speech data to generate a speech translation output. As another example, the machine-learned model(s) can process the speech data to generate a latent embedding output. As another example, the machine-learned model(s) can process the speech data to generate an encoded speech output (e.g., an encoded and / or condensed representation of the speech data, etc.). As another example, the machine-learned model(s) can process the speech data to generate an upscaled speech output (e.g., speech data of higher quality than the input speech data, etc.). As another example, the machine-learned model(s) can process the speech data to generate a text representation output (e.g., a text representation of the input speech data, etc.). As another example, machine-learned model(s) can process speech data to generate predicted outputs.

[0043] In some implementations, input to the machine-learned model(s) of the present disclosure can be latent-coded data (e.g., a latent space representation of the input, etc.). The machine-learned model(s) can process the latent-coded data to generate an output. As an example, the machine-learned model(s) can process the latent-coded data to generate a recognition output. As another example, the machine-learned model(s) can process the latent-coded data to generate a reconstruction output. As another example, the machine-learned model(s) can process the latent-coded data to generate a retrieval output. As another example, the machine-learned model(s) can process the latent-coded data to generate a reclustering output. As another example, the machine-learned model(s) can process the latent-coded data to generate a prediction output.

[0044] In some implementations, input to the machine-learned model(s) of the present disclosure can be statistical data. The statistical data can be, represent, or otherwise include data calculated and / or computed from some other data source. The machine-learned model(s) can process the statistical data to generate an output. As an example, the machine-learned model(s) can process the statistical data to generate a recognition output. As another example, the machine-learned model(s) can process the statistical data to generate a prediction output. As another example, the machine-learned model(s) can process the statistical data to generate a classification output. As another example, the machine-learned model(s) can process the statistical data to generate a segmentation output. As another example, the machine-learned model(s) can process the statistical data to generate a visualization output. As another example, the machine-learned model(s) can process the statistical data to generate a diagnostic output.

[0045] In some implementations, input to the machine-learned model(s) of the present disclosure can be sensor data. The machine-learned model(s) can process the sensor data to generate an output. As an example, the machine-learned model(s) can process the sensor data to generate a recognition output. As another example, the machine-learned model(s) can process the sensor data to generate a prediction output. As another example, the machine-learned model(s) can process the sensor data to generate a classification output. As another example, the machine-learned model(s) can process the sensor data to generate a segmentation output. As another example, the machine-learned model(s) can process the sensor data to generate a visualization output. As another example, the machine-learned model(s) can process the sensor data to generate a diagnostic output. As another example, the machine-learned model(s) can process the sensor data to generate a detection output.

[0046] In some cases, the machine-learned model(s) may be configured to perform a task that includes encoding input data for reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task may be an audio compression task. The input may include audio data, and the output may include compressed audio data. In another example, the input includes visual data (e.g., one or more images or videos), and the output includes compressed visual data, and the task is a visual data compression task. In another example, the task may include generating an embedding for the input data (e.g., input audio or visual data).

[0047] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data of one or more images and the task is an image processing task. For example, the image processing task can be image classification, and the output is a set of scores, each score corresponding to a different object class and representing the likelihood that one or more images depict an object belonging to the object class. The image processing task can be object detection, and the image processing output identifies one or more regions in one or more images and, for each region, the likelihood that the region depicts an object of interest. As another example, the image processing task can be image segmentation, and the image processing output defines, for each pixel in one or more images, a respective likelihood for each category in a predetermined category set. For example, the category set can be foreground and background. As another example, the category set can be object classes. As another example, the image processing task can be depth estimation, and the image processing output defines, for each pixel in one or more images, a respective depth value. As another example, the image processing task can be motion estimation, where the network input includes multiple images and the image processing output defines, for each pixel of one of the input images, the motion of the scene depicted in pixels between the images in the network input.

[0048] In some cases, the input includes audio data representing a speech utterance and the task is a speech recognition task. The output may include text output that is mapped to the speech utterance. In some cases, the task includes encrypting or decrypting input data. In some cases, the task includes a microprocessor performance task, such as branch prediction or memory address translation.

[0049] 1A illustrates one exemplary computing system that can be used to implement the present disclosure. Other computing systems can be used as well. For example, in some implementations, a user computing device 102 can include a model trainer 160 and a training dataset 162. In such implementations, the model 120 can be trained and used locally on the user computing device 102. In some such implementations, the user computing device 102 can implement a model trainer 160 that personalizes the model 120 based on user-specific data.

[0050] 1B illustrates a block diagram of an exemplary computing device 10 that performs content generation within a predictive content generation space, according to some implementations of the present disclosure. The computing device 10 can be a user computing device or a server computing device.

[0051] Computing device 10 includes several applications (e.g., applications 1-N). Each application includes its own machine learning library and machine-learned model(s). For example, each application may include a machine-learned model. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.

[0052] 1B , each application may communicate with several other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application may communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.

[0053] 1C illustrates a block diagram of an example computing device 50 for training machine-learned model(s) to generate predictive content, according to some implementations of the present disclosure. The computing device 50 can be a user computing device or a server computing device.

[0054] Computing device 50 includes several applications (e.g., applications 1-N). Each application communicates with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and the model(s) stored therein) using an API (e.g., a common API across all applications).

[0055] The central intelligence layer includes several machine-learned models. For example, as illustrated in FIG. 1C , each machine-learned model may be provided for each application and managed by the central intelligence layer. In other embodiments, two or more applications may share a single machine-learned model. For example, in some implementations, the central intelligence layer may provide a single model for all applications. In some implementations, the central intelligence layer is included within or otherwise implemented by the operating system of the computing device 50.

[0056] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for the computing device 50. As illustrated in FIG. 1C , the central device data layer can communicate with several other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0057] FIG. 2A illustrates an exemplary layout of an interface 200 for a predictive content generation space at a first time, according to some embodiments of the present disclosure. Specifically, as shown, the interface 200 for a predictive content generation space is a two-dimensional interface 200 including a background 201 and a toolbar 204 including tools 204A, 204B, 204C, 204D, and 204E. The background of the interface 200 is a solid background, such as a whiteboard background. Note that FIGS. 2A-4B depict only a two-dimensional whiteboard to more clearly illustrate various aspects of the present disclosure. However, embodiments of the present disclosure are not limited to either the background or the interface. For example, the background 202 of the interface 200 may instead depict a traditional chalkboard background. In other examples, the background 202 of the interface 200 may depict a background selected by the user.

[0058] Additionally or alternatively, in some implementations, interface 200 may be overlaid on the interface of a separate application. For example, the predictive content generation space may be an operating system (OS)-level feature that can run concurrently with other applications. A user may browse Internet content via a web browser and then run the predictive content generation space. The background 202 of interface 200 may be the web page the user is browsing.

[0059] Interface 200 may include a toolbar 204 that provides access to tools 204A-204E, which are brush tools that a user can utilize to create (i.e., draw) lines, shapes, points, etc. within predictive content generation space interface 200. As shown in Figure 5, each of tools 204A-204E may be associated with one of several machine learning tasks.

[0060] Turning to FIG. 5, FIG. 5 illustrates a data structure 500 associating tools 204A-204E with machine learning tasks, according to some implementations of the present disclosure. For example, tool 204A can be associated with a content expansion task 502A. Tool 204B can be associated with a content atomization task 502B. Tool 204C can be associated with a content analysis task 502C. Tool 204D can be associated with a prompt generation task 502D. Tool 204E can be associated with a content synthesis task 502E. In some implementations, each machine learning task 502 can be associated with model processing instructions 504. The model processing instructions 504 can indicate a series of processing steps required to complete the corresponding machine learning task 502. For example, for an input image with content expansion task 502A, the corresponding model processing instructions 504 indicate that the image should be processed with a machine-learned semantic processing model to obtain intermediate output, which may then be processed with a machine-learned search model to obtain predicted content.

[0061] In some implementations, tools 204A-204E can be associated with user history data 506. User history data 506 can describe a particular user's previous use of tools 204A-204E. For example, user history data 506 may indicate that a user rarely uses tool 204D but commonly uses tool 204A. Thus, when automatically selecting tools 204A-204E for a user, a computing system (e.g., server computing system 130, user computing device 102, etc.) can determine which of tools 204A-204E to use based on user history data 506.

[0062] 2A , while tools 204A-204E are depicted as brush tools, it should be noted that tools 204A-204E are not necessarily limited to implementations of brush tools. Rather, toolbar 204 may include tools in any manner operable to select content elements 206 within predictive content generation space interface 200. For example, tool 204A may instead be a box tool that a user can use to surround a content item of interest. In another example, tool 204A may instead be a writing tool that allows a user to write instructions directly into predictive content generation space interface 200.

[0063] The predictive content generation space interface 200 may include one or more content elements 206. As shown, the content element 206 may include multiple portions 206A, 206B, 206C, and 206D. Each of these portions may depict one or more entities within the content element. As previously mentioned, the content element 206 may include any type or combination of multimedia content (e.g., audio data, text content, URLs, video data, image(s), three-dimensional representation(s), video games, web pages, summaries, live streams, etc.).

[0064] At a first time T1, a user may utilize a tool on toolbar 204 to select at least a portion of content element 206 depicted within predictive content generation space interface 200. As shown, the user uses tool 204B to provide selection input 208 that intersects with portion 206D of content element 206. Specifically, selection input 208 is a line that begins at location 208A, intersects with portion 206D, and terminates at location 208B. In some implementations, intersecting selection input 208 with portion 206D allows the user to select portion 206D of content element 206 while excluding portions 206A-206C. Alternatively, in some implementations, intersecting selection input 208 with portion 206D may select the entire content element 206 (e.g., all portions 206A-206D).

[0065] Note that while the predictive content generation space is depicted as a two-dimensional space, it is not limited to two-dimensional space. For example, the predictive content generation space may be a virtual three-dimensional space, and interface 200 may be displayed within a display device of an augmented reality (AR) / virtual reality (VR) device. In some implementations, toolbar 204 may include tools for a three-dimensional environment. For example, tool 204A may be a wand that allows a user to draw in three dimensions. In other examples, tool 204A may simulate interaction between a user's attachment and an augmented reality element (e.g., a three-dimensional rendering, etc.), allowing a user to directly interact with the content element with their hand as an augmented reality element. For example, the content element may be a three-dimensional box showing a video from a video sharing site. The user can provide a selection input by touching the content element. In response, the video may be played within the box, or alternatively, may be played directly to the user via a different interface (e.g., two-dimensional interface 200, etc.). In this manner, the interface 200 may switch between a two-dimensional interface and a three-dimensional interface to facilitate user interaction with content elements using the tools in the toolbar 204.

[0066] Although not shown, it should be noted that a user can provide content element 206 directly to interface 200 for a predictive content generation space. For example, content element 206 can be an image, and a user can “drag and drop” content element 206 from a file storage system onto interface 200 for a predictive content generation space to begin uploading the image to interface 200. In another example, a user may enter text content directly into interface 200 in a free-form manner (e.g., by clicking and typing in a location on interface 200, a “text box” can be created in which a content element containing text content can be created). In yet another example, a user can copy a URL to a video hosted on a hosting website and paste the URL into interface 200. The video can then be displayed directly within content element 206. Accordingly, it should be broadly understood that interface 200 can be configured to enable a user to “drop,” upload, or otherwise provide content of any type or form (e.g., video, images, search queries, multimodal search queries, video games, AR / VR objects, URLs, web applications, virtual compute instances, etc.) into interface 200, which can then be displayed within a content element (e.g., content element 206) of interface 200.

[0067] For example, in some implementations, the content element 206 may be a query (e.g., a query image, a text query, a voice utterance containing a query, etc.). For example, the interface 200 for a predictive content generation space may allow a user to generate text content directly within the predictive content generation space. The user can then select a content element (e.g., a query) and perform a search.

[0068] Additionally or alternatively, in some implementations, content element 206 can include multiple content elements that collectively form a multimodal search query. For example, a user may enter a text query for "blue shoes" in a text box content element depicted in interface 200. The user may then "drag and drop" an image of a white shoe within interface 200 to form a content element that includes the image of the white shoe. The user can select both the text box content element and the image content element to provide a multimodal search query that can be processed by a machine-learned model.

[0069] FIG. 2B illustrates an example interface 200 for machine-learned generation of predictive content elements within a predictive content generation space, according to some implementations of the present disclosure. Specifically, a portion 206D of a content element 206 selected by a user via a selection input 208 can be processed by machine-learned model(s) 209 (e.g., machine-learned model(s) 120 or 140 of FIG. 1A ) to obtain predictive content 210. The machine-learned model(s) 209 can be machine-learned model(s) trained to perform a machine-learning task associated with tool 204B. For example, the data description of portion 206D can be image data. Tool 204B can be associated with a machine-learning task for content development (e.g., searching for content similar to input content). The machine-learned model(s) 209 can be or otherwise include a semantic image processing model trained to process an input image and obtain (or generate) semantically similar images. The predictive content 210 may include the captured / generated image, or alternatively may include data indicating the image (e.g., a hyperlink to where the image is hosted, a pointer to where the image is stored in memory, etc.).

[0070] Note that the machine learning task for which machine-learned model(s) 209 are trained may be one in which the output includes multiple types of media. Following the previous example, machine-learned model(s) 209 may include other models (or the same models) configured to retrieve information about data describing images included in portion 206D. For example, the semantic image processing model of the previous example may process data describing images in portion 206D to obtain a semantic description output. One model may use the semantic description output to obtain similar images (or may obtain images in a conventional manner), while another model may generate or obtain information about the semantic description output. For example, portion 206D may depict a duck swimming in a pond. The semantic description output may indicate that the image depicts a duck swimming in a pond. Another model (e.g., a large-scale language model, etc.) in machine-learned model 209 may generate information about the duck's history.

[0071] Based on the predictive content 210, the predictive content element generator 211 can generate predictive content elements 212A-212D. The predictive content elements 212A-212D content elements describe the predictive content 210. For example, the predictive content 210 may include an image. The predictive content element 212A may depict or otherwise include an image. In another example, the predictive content 210 may include textual content. The predictive content element 212A may include textual content, a summary of the textual content, or a link to a location where the textual content is hosted. In yet another example, the predictive content 210 may include a cloud-based video game. The predictive content element 210A may include being configured to execute the cloud-based video game when the predictive content element 212A is selected by a user.

[0072] It should be noted that the operations of the machine-learned model(s) 209 and the predictive content element generator 211 are depicted as occurring outside of the interface 200 solely to illustrate that the operations are not depicted within the interface 200. Thus, the depicted locations of the machine-learned model(s) 209 and the predictive content element generator 211 should not be interpreted as indicating which computing device(s) are utilized to perform the operations of the machine-learned model(s) 209 and the predictive content element generator 211.

[0073] 2C shows an example layout of interface 200 for a predictive content generation space at a second time, according to some embodiments of the present disclosure. Specifically, at time T2, predictive content elements 212A, 212B, 212C, and 212D are generated and depicted within interface 200 of the predictive content generation space. In some embodiments, connection interface element 212 can be generated within the predictive content generation space. Connection interface element 212 can depict connections between predictive content elements 212A-212D and portion 206D of content element 206.

[0074] Note that in some implementations, the predictive content elements 212A-212D may be generated and rendered at a location within the interface 200 where the user completed the selection input 208. For example, as shown in FIG. 2A, the user completed the selection input 208 at location 208B. Thus, the predictive content elements 212A-212D are generated and rendered at approximately the same location as location 208B. Alternatively, in some implementations, the location at which the predictive content elements 212A-212D are generated and rendered may be determined in some other manner (e.g., based on user preferences, user history data, the type of content of the predictive content elements 212A-212D, etc.).

[0075] At time T2, portion 206D of content element 206 selected by user selection input 208 is processed with a machine-learned model. More specifically, data describing portion 206D is processed with the machine-learned model to obtain predictive content. The machine-learned model can be a model trained to perform machine learning tasks associated with tool 204B. Predictive content elements 210A, 210B, 210C, and 210D can be generated based on the predictive content.

[0076] 2D illustrates content elements and predicted content elements within interface 200 of a predictive content generation space at a second time, according to some implementations of the present disclosure. Specifically, note that FIG. 2C merely presents an alternative view of FIG. 2C illustrating the content depicted within content elements 206 and 212A-212D, rather than the layout of content elements 206 and 212A-212D. For example, content element 206 includes an image depicting the overhead layout of a room. The room includes a sofa, a television, a table, a plant, and a ping-pong table.

[0077] As shown, the plant is located within portion 206D of content element 206, which the user selected with selection input 208 using tool 204B. Accordingly, predictive content elements 212A-212D include content corresponding to the machine learning task associated with tool 204B. In this example, tool 204B can be a content extraction tool, whose task is to find content similar to the content of portion 206D. Continuing with this example, the plant depicted in portion 206D can be a sunflower. Predictive content 210 obtained or generated using machine-learned model 209 can relate to sunflowers. For example, content element 212A can include information about the sunflower plant (e.g., information obtained from an online dictionary, information synthesized using a large-scale language model, etc.). Predictive content element 212B can include an image and a link to a video hosted on a video sharing site related to how to grow sunflowers. Predictive content element 212C can be an image of a sunflower. The predictive content element 212D can be a classification of concept tags related to room decor (i.e., the purpose of sunflowers). For example, if a user selects a furniture content tag from the predictive content element 212D (e.g., by touch input, click input, etc.), a second content element can be generated that includes predictive content related to furniture. This furniture may be related to furniture depicted in other portions of the content element 206.

[0078] 2D further illustrates data 214 describing the content elements 206. Specifically, the data describing the content elements 206 is metadata 214 about images that collectively describe an overhead view of a room. For example, the metadata 214 describes each entity depicted within the room (e.g., a television, a sofa, a table, a plant, etc.). Furthermore, the metadata 214 can describe a semantic view of the image (e.g., the image shows an overhead view of a family room). In some implementations, the metadata may already be included in the content elements 206. For example, the content elements may include images that describe the room, and the images may include the metadata 214. Alternatively, in some implementations, the metadata 214 may be determined. For example, the content elements 206 may be processed with machine-learned model(s) 209 to determine the metadata 214.

[0079] It should be noted that while the metadata 214 describes the entire content element 206, it may instead describe only the relevant portion of the content element 206. Following the previous example, after the user selects portion 206D with the selection input 208, the metadata 214 may be determined for portion 206D.

[0080] 3A illustrates an exemplary layout of an interface 200 in which an alternative user selection input is provided a first time, according to some other implementations of the present disclosure. Specifically, as shown in FIG. 2A , the selection input 208 from the user may be a line that intersects with a portion of the content element 206. However, the user is not limited to such selection inputs, nor is the user limited to selecting a portion(s) of the content element 206. For example, as shown, the user may provide a selection input 302 that is a closed shape input. Specifically, the user can use the brush tool 204B to draw a closed shape around the content element 206 to provide a selection input 302 that selects the entire content element 206. In this manner, the user can indicate interest in the entire content element 206, rather than a specific portion.

[0081] Alternatively, in some implementations, the user may provide a click input 303. Specifically, the user may select the brush tool 204B and then draw a "point" (i.e., click or touch a location on the interface 200) to provide a selection input 302 that selects a portion or the entire content element 206. For example, the user may click toward the center of the content element 206 to indicate interest in the entire content element 206. In other examples, the user may click to provide a selection input 303 that is away from the center of the content element 206 to indicate interest in a particular portion of the content element 206. In some implementations, the user's intent associated with the selection input 303 may be determined based on user history data and / or the content of the content element 206.

[0082] 3B illustrates predictive content elements corresponding to a selection of the entire content element 206 in interface 200 at a second time, according to some embodiments of the present disclosure. Specifically, at time T2, after the entire content element 206, predictive content elements 304A, 304B, 304C, and 304D are generated and rendered in interface 200 as described with respect to predictive content elements 212A-212D.

[0083] For example, predictive content element 304A may include a summary of a website showing whether programming related to a home is available for streaming. Predictive content element 304B may include an image of a table tennis racket that is popular among professional players. Predictive content element 304C may include the same image of a sunflower as predictive content element 212C of FIG. 2D. Predictive content element 304D may include the same taxonomy of concept tags related to room decor as predictive content element 212D of FIG. 2D. FIG. 3B also shows a second selection input 306 made by a user using tool 204B. Further, the user provides second selection input 306 using brush tool 204B. The second selection input 306 selects predictive content element 304B, which depicts a table tennis racket that is popular among professional players.

[0084] FIG. 3C illustrates predictive content elements corresponding to the selection of predictive content element 304B within interface 200 at a third time, according to some implementations of the present disclosure. Specifically, at time T3, second content elements 308A, 308B, and 308C can be generated and depicted within interface 200 as described with respect to FIG. 2B. Because selection input 306 selected content element 304B with content deployment tool 204B depicting an image of a table tennis racket, second content elements 308A-308C can be related to table tennis. For example, second predictive content element 308A, when selected by a user, can include a prompt that can search for table tennis coaches in the user's area. Second predictive content element 308B is a video from a video hosting site related to a table tennis tournament. Second predictive content element 308C is an article related to table tennis equipment.

[0085] In some implementations, additional connection elements 212 can be generated to link second predictive content elements 308A-308C to content element 304B. In this manner, the predictive content generation space can be utilized as a continuous surface for users to creatively explore content and ideas within a single space.

[0086] 4A illustrates an example layout of interface 200 in which an audio brush is selected with an audio utterance provided by a user at a first time, according to some other implementations of the present disclosure. Specifically, rather than selecting tool 204B, the user may instead select tool 204C. Tool 204C may correspond to an audio brush tool. The audio brush tool may be configured to select content element(s) or portion(s) of a content element in the same manner as brushes 204A-204B and 204D-204E. However, the machine learning task associated with audio brush tool 204C may be specified by an audio utterance 404 provided by the user.

[0087] For example, a user can provide a selection input 402 that selects the entire content element 206 using the voice brush tool 204C. At the same time, the user can also provide a voice utterance 404 that indicates an interest in what type of television is depicted within the content element 206. In this manner, the user can indicate a particular portion or entity of interest within the content element 206 via the voice utterance 404. Additionally, the user can indicate which machine learning task should be performed via the voice utterance 404. For example, by asking, "What kind of TV is this?" the user can indicate that a content analysis task is desired.

[0088] In some implementations, the machine learning task indicated by audio utterance 404 can be a machine learning task associated with another tool in toolbar 204. Following the previous example, a content analysis machine learning task may be associated with tool 204E. However, indicating the content analysis task via audio utterance 404 allows the content analysis task to be performed without the user having to manually select tool 204E.

[0089] In some implementations, the machine learning task indicated by the speech utterance 404 may not explicitly match an existing machine learning task. For example, a user may simply indicate an interest in television without a corresponding machine learning task (e.g., "this is good television"). In response, the large-scale language model of the machine-learned model(s) 209 may be utilized to determine the intended or optimal machine learning task.

[0090] In some implementations, a machine learning task may be selected based on user history data describing a user's previous interactions within the predictive content generation space. For example, the user history data may indicate that in a previous interaction, the user used the content development tool 204B to provide a selection input that is similar to the selection input 402. Accordingly, a content development task may be selected for assignment to the audio brush tool 204C.

[0091] Additionally or alternatively, in some implementations, a machine learning task may be selected based on data describing at least a portion of the content element. For example, metadata 214 of FIG. 2D may indicate that content element 206 includes an image depicting a television. User history data may indicate that a user previously selected an image of a television using content development tool 204B. Alternatively, generalized user history data may indicate that most users who select an image of a television do so with content development tool 204B. In this manner, a machine learning task may be assigned to the voice brush tool regardless of whether the task is explicitly indicated in voice utterance 404.

[0092] 4B illustrates predicted content elements corresponding to the selection of content element 206 in interface 200 via the audio brush tool at a second time, according to some implementations of the present disclosure. Specifically, as shown, content elements 406A, 406B, and 406C can be obtained / generated and depicted in interface 200 according to audio brush tool 204C and audio utterance 404. In this example, a content development task is assigned to audio brush tool 204C according to audio utterance 404. For example, predicted content element 406A can include information about marketplaces where televisions can be purchased. Predicted content element 406B can be information obtained from an online dictionary and summarized in the content element. Predicted content element 406C can be a summary or highlight of information obtained from an online forum.

[0093] While only a single user is shown selecting content element(s) in this example, it should be noted that embodiments of the present disclosure are not limited to a single user. Rather, the predictive content generation space can be utilized by multiple users. For example, the predictive content generation space may be a collective space in which multiple users can independently interact. For example, a first user may generate a "web" of content elements at a first location within the predictive content generation space and then navigate to a second location in the predictive content generation space to generate a new "web" of content elements. At a later time, a second user may explore the predictive content generation space and discover the "web" of content elements at the first location by the first user. The second user may then continue to deploy the content elements at the first location or may navigate to a different location to create a new web of content elements.

[0094] As another example, multiple users may interact with content elements simultaneously. For example, two users may use two different brushes to select content elements in the predictive content generation space (e.g., one user "brushes" to the right, generating additional content elements on the right, while the other user "brushes" to the left, generating additional content elements on the left). In this manner, multiple users can collectively create a web of content elements as they explore predictive content in a collaborative manner.

[0095] As yet another example, the state of a predictive content generation space can be saved and distributed to other users of the predictive content generation space. For example, a user may generate content within an instance of a predictive content generation space. The user can save the state of the predictive content generation space. The state of the predictive content generation space can include any content elements generated within the space, the location of the predictive content within the space, and any connection elements that exist to link the content elements within the space. The saved state of the predictive content generation space can be provided to several other users who have run separate instances of the predictive content generation space. The other users can load the saved state of the predictive content generation space to obtain the saved content elements, locations, connection elements, etc. In this manner, a session in which a user creates and explores predictive content within a predictive content generation space can be shared with other users.

[0096] Model layout example 6 illustrates a block diagram of an example machine-learned model 600 (e.g., machine-learned model(s) 209, machine-learned model(s) 120, machine-learned model(s) 140, etc.) according to some implementations of the present disclosure. In some implementations, machine-learned model 600 is trained to receive a set of input data 604 that describes at least a portion of content elements and, as a result of receiving input data 604, provide output data 606 that is, or is otherwise indicative of, predicted content.

[0097] In some implementations, the machine-learned model 600 can include a task-specific portion 602 operable to perform a particular machine-learned task. For example, the machine-learned model 600 can be a large-scale language model trained to perform multiple machine-learned tasks. The task-specific portion 602 can be a portion trained to perform a single task. For example, the task-specific portion 602 can be trained to generate summaries of information for inclusion in content elements.

[0098] 7 shows a block diagram of an example of a machine-learned model 700 according to some other embodiments of the present disclosure. The machine-learned model 700 is similar to the machine-learned model 600 of FIG. 6, except that the machine-learned model 700 further includes a preprocessing model 702. The preprocessing model can be any model trained to process input data 604 and generate intermediate outputs 704 that can be processed by the task-specific model 602.

[0099] Continuing with the previous example, the task-specific portion 602 can be configured to summarize the obtained information for inclusion in a content element of the output data 606. The input data 604 can be data describing the content element, including an image. The preprocessing model can be a model trained to semantically process the image and output an intermediate output 704 that includes a semantic description of the image. The intermediate output 704 (i.e., the semantic description of the image) can be processed by the task-specific model 602 to obtain the output data 606 (e.g., predicted content including a summary).

[0100] 8 shows a block diagram of an example of a machine-learned model 800 according to some other embodiments of the present disclosure. The machine-learned model 800 is similar to the machine-learned model 700 of FIG. 7, except that the machine-learned model 800 further includes a task selection model 803. The task selection model 803 can be a model trained to process context information 804 and output a selection of a machine-learning task.

[0101] For example, context information 804 may include user history data describing users of the predictive content generation space, as described above. Additionally, in some implementations, context information 804 may describe the context in which the predictive content generation space is being used (e.g., whether it is being used collaboratively, the subject matter based on all content elements within the space, etc.).

[0102] In some implementations, task selection model 803 may process context information 804 to obtain task selection output 806. Additionally, in some implementations, task selection model may also process input data 604 describing the selected content elements. Task selection output 806 may be provided to preprocessing model 702 and task-specific model 602. Continuing with the previous example, task selection model 803 may process context information 804 indicating that a user prefers to utilize a content development tool. Task selection model 803 may also process input data 604 indicating that the content elements include image data (e.g., image data may have a strong correlation with content development tasks).

[0103] The task selection model 803 can provide a task selection output 803 to the preprocessing model 702 and the task-specific model 602. For example, the preprocessing model 702 can be a large model configured to be trained to perform multiple preprocessing tasks. Based on the task selection output 806, one of the multiple preprocessing tasks can be indicated to the preprocessing model 702. Similarly, the task-specific model 602 can be one of several task-specific models or a large model trained to perform multiple tasks. Based on the task selection output 806, a model can be selected as the task-specific model 602, or a task can be indicated to the task-specific model 602. In this manner, multiple conventional large-scale machine-learned models can be utilized to perform multiple processing operations that facilitate the predictive content generation space.

[0104] Exemplary Methods 9 illustrates a flowchart diagram of an exemplary method for generating content within a predictive content generation space, according to an exemplary embodiment of the present disclosure. While FIG. 9 illustrates steps occurring in a particular order for purposes of illustration and explanation, the methods of the present disclosure are not limited to the specifically depicted order or arrangement. Various steps of method 900 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0105] At 902, the computing system uses a first tool of a plurality of tools to obtain data indicative of a selection of content elements to be depicted within the predictive content generation space. Specifically, the computing system uses a first tool of a plurality of tools of the predictive content generation space to obtain data indicative of a selection by a user of at least a portion of the content elements to be depicted within the predictive content generation space. The plurality of tools can each be associated with a plurality of machine learning tasks. Each of the plurality of tools can be operable to select at least a portion of each of one or more content elements to be depicted within the predictive content generation space.

[0106] In some implementations, the first machine learning task can include a content development task, where the machine-learned model is trained to process data describing at least a portion of a content element and output content that is similar to at least a portion of the content element.

[0107] In some embodiments, the first machine learning task includes a content analysis task, and the machine learning model is trained to process data describing at least a portion of the content elements and output a summary of at least a portion of the content elements.

[0108] In some embodiments, the first machine learning task includes a prompt generation task, and the machine learning model is trained to process data describing at least a portion of the content element and output one or more prompts to the user related to at least a portion of an aspect of the content element.

[0109] In some implementations, obtaining data indicative of a user selection of at least a portion of the content elements further includes selecting a machine learning task from a plurality of machine learning tasks. The task can be selected based at least in part on historical user data describing a user's previous interactions within the predictive content generation space and / or data describing at least a portion of the content elements. The computing system can assign the machine learning task to a first tool.

[0110] In some implementations, the content elements include one or more of images, video data, three-dimensional displays, text content, URLs, audio data, video games, and the like.

[0111] In some implementations, to obtain data indicative of a user selection of at least a portion of the content element, the computing system may obtain data indicative of a user selection of a first portion of an image, where the first portion of the image may depict a first entity and the second portion of the image may depict a second entity different from the first entity.

[0112] In some implementations, prior to processing the data describing at least a portion of the content element with the machine-learned model, the computing system may determine the data describing at least a portion of the content element. In some implementations, the data describing at least a portion of the content element may include metadata associated with the content element.

[0113] In some embodiments, the predictive content generation space is a two-dimensional space. In some embodiments, the predictive content generation space is displayed on top of a separate application interface. In some embodiments, the predictive content generation space is a three-dimensional augmented reality (AR) / virtual reality (VR) space.

[0114] In some implementations, the data indicating the selection of a content element may be or otherwise include data indicating a multimodal search query. For example, the content element selected by the user may be a text query. The user may also select a content element that is an image simultaneously with selecting the text query.

[0115] At 904, the computing system processes data describing the content elements with the machine-learned models to obtain predictive content. Specifically, the computing system processes data describing at least a portion of the content elements with the machine-learned models to obtain predictive content. The machine-learned models can be trained to perform first machine learning tasks respectively associated with the first tools. In some implementations, the predictive content can include predictive content for a first content type corresponding to the first machine learning task.

[0116] In some implementations, the first machine learning task can be a multimodal search task. Following the previous example, a user may select multiple content elements (e.g., select a text query, image, and video data, etc.) to form a multimodal search query. A machine-learned model can be trained to retrieve search results based on the multimodal query or to facilitate retrieval of multimodal results (e.g., generate embeddings that can be used to retrieve results from a multimodal search space).

[0117] In some embodiments, processing the data describing at least a portion of the content element includes processing the data describing the first portion of the content element with a machine-learned model to obtain predictive content that includes information identifying one or more images that are semantically similar to the first portion of the image of the content element.

[0118] In some implementations, the machine-learned model may include multiple machine-learned models that collectively process inputs in an order specified by the corresponding machine-learning tasks.

[0119] In some embodiments,

[0120] At 906, the computing system generates one or more predictive content elements within the predictive content generation space. The one or more predictive content elements can describe the predictive content.

[0121] In some implementations, the first machine learning task can be a machine-learned semantic image retrieval task, and the content element can include an image. To process the data describing at least a portion of the content element, the computing system can process the data describing at least a portion of the content element with the machine-learned model to obtain predictive content including information identifying one or more images that are semantically similar to the image of the content element. To generate the one or more predictive content elements, the computing system can generate one or more predictive content elements within a predictive content generation space, each including one or more images.

[0122] In some implementations, a computing system can obtain data indicating a user's selection of one of the one or more predictive content elements depicted within the predictive content generation space using a second tool from the plurality of tools, the second tool being different from the first tool. The computing system can process the data describing the predictive content element with a machine-learned model to obtain the second predictive content. The machine-learned model can be trained to perform second machine learning tasks respectively associated with the second tool. The computing system can generate one or more second predictive content elements within the predictive content generation space. The one or more second predictive content elements can describe the second predictive content.

[0123] In some implementations, the computing system can generate connection elements within the predictive content generation space that depict connections between the predictive content element and one or more second predictive content elements.

[0124] In some implementations, the machine-learned model includes a large-scale language model that is trained to perform both the first machine learning task and the second machine learning task.

[0125] In some implementations, the first tool includes a brush tool. Obtaining data indicative of a user selection of at least a portion of the content element can include obtaining data indicative of a shape generated by the user using the brush tool within the predictive content generation space and determining the user-generated shape that selects at least a portion of the content element to be depicted within the predictive content generation space.

[0126] In some implementations, the user-generated shape includes a line. Determining that the user-generated shape selects at least a portion of the content element can include determining that the user-generated line intersects at least a portion of the content element depicted in the predictive content generation space.

[0127] In some implementations, the user-generated shape includes a closed shape. Determining that the user-generated shape selects at least a portion of the content element can include determining that the user-generated closed shape includes at least a portion of the content element depicted in the predictive content generation space.

[0128] In some implementations, the user-generated shape includes a point corresponding to a touch input or a click input. Determining that the user-generated shape selects at least a portion of the content element can include determining that the user-generated point is located on at least a portion of the content element depicted within the predictive content generation space.

[0129] In some implementations, the second tool can include an audio brush tool, and the second machine-learned task can include a speech recognition task. Obtaining data indicative of a user's selection of a predictive content element using the second tool can include obtaining data indicative of a line generated by the user within the predictive content generation space using the audio brush tool. The computing system can determine that a line generated by the user using the audio brush tool intersects with a predictive content element depicted within the predictive content generation space. The computing system can obtain data describing an audio utterance by the user. The audio utterance can indicate a third tool of the plurality of tools, different from the first tool and the second tool. The third tool can be associated with a third machine learning task of the plurality of machine learning tasks. Processing the data describing the predictive content element with the machine-learned model can include processing the data describing the audio utterance with the machine-learned model trained to perform the second machine learning task to obtain a speech recognition output that identifies the third tool. Based on the speech recognition output, the computing system can process the data describing the predictive content elements using a machine-learned model trained to perform a third machine learning task to obtain second predictive content.

[0130] Additional Disclosures The technology described herein refers to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functionality among components. For example, the processes described herein can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0131] While the present subject matter has been described in detail with reference to various specific exemplary embodiments thereof, each example is provided for purposes of illustration and not as a limitation of the present disclosure. Those skilled in the art, upon understanding the foregoing, may readily create modifications, variations, and equivalents to such embodiments. Accordingly, the present disclosure does not exclude the inclusion of such modifications, variations, and / or additions to the present subject matter that would be readily apparent to those skilled in the art. For example, features illustrated or described as part of one embodiment may be used with other embodiments to create yet another embodiment. Accordingly, the present disclosure is intended to cover such modifications, variations, and equivalents.

Claims

1. 1. A computer-implemented method for generating content within a predictive content generation space via user-specified machine learning tasks, comprising: obtaining, by a computing system including one or more computing devices, data indicative of a user's selection of at least a portion of a content element depicted within a predictive content generation space using a first tool of a plurality of tools of the predictive content generation space; the plurality of tools are respectively associated with a plurality of machine learning tasks; each of the plurality of tools is operable to select at least a portion of each of one or more content elements depicted within the predictive content generation space; To get and processing, by the computing system, data describing at least a portion of the content elements with machine-learned models to obtain predictive content, the machine-learned models each being trained to perform a first machine-learning task associated with the first tool; generating, by the computing system, one or more predictive content elements within the predictive content generation space, the one or more predictive content elements describing the predictive content; A computer-implemented method comprising:

2. The computer-implemented method of claim 1 , wherein the predictive content comprises predictive content of a first content type corresponding to the first machine learning task.

3. the first machine learning task is a machine-learned semantic image retrieval task; the content elements include images; processing the data describing the at least a portion of the content element includes processing, by the computing system, the data describing the at least a portion of the content element with the machine-learned model to obtain predictive content, the predictive content including information identifying one or more images that are semantically similar to the image of the content element; 3. The computer-implemented method of claim 2, wherein generating the one or more predictive content elements comprises generating, by the computing system, one or more predictive content elements within the predictive content generation space that each include the one or more images.

4. obtaining the data indicative of a selection by the user of the at least a portion of the content element includes obtaining, by the computing system, data indicative of a selection by the user of a first portion of the image, the first portion of the image depicting a first entity and a second portion of the image depicting a second entity different from the first entity; 4. The computer-implemented method of claim 3, wherein processing the data describing the at least a portion of the content element includes processing, by the computing system, the data describing the first portion of the content element with the machine-learned model to obtain predictive content including information identifying one or more images that are semantically similar to the first portion of the image of the content element.

5. obtaining, by the computing system, data indicative of a selection by the user of one of the one or more predictive content elements depicted within the predictive content generation space using a second tool of the plurality of tools, the second tool being different from the first tool; processing, by the computing system, data describing the predictive content elements with machine-learned models to obtain second predictive content, the machine-learned models each being trained to perform second machine learning tasks associated with the second tool; generating, by the computing system, one or more second predictive content elements within the predictive content generation space, the one or more second predictive content elements describing the second predictive content; The computer-implemented method of any one of claims 1 to 4, further comprising:

6. 6. The computer-implemented method of claim 5, wherein the method further includes generating, by the computing system, a connection element within the predictive content generation space that depicts a connection between the predictive content element and the one or more second predictive content elements.

7. 7. The computer-implemented method of claim 5, wherein the machine-learned model comprises a large-scale language model trained to perform both the first machine learning task and the second machine learning task.

8. the first tool includes a brush tool; Obtaining the data indicative of the selection by the user of the at least a portion of the content element comprises: obtaining, by the computing system, data indicative of a shape generated by the user within the predictive content generation space using the brush tool; determining, by the computing system, that the shape generated by the user selects at least a portion of the content elements depicted within the predictive content generation space; 8. The computer-implemented method of any one of claims 1 to 7, comprising:

9. the user-generated shape includes a line; 9. The computer-implemented method of claim 8, wherein determining that the shape generated by the user selects the at least a portion of the content element includes determining, by the computing system, that the line generated by the user intersects the at least a portion of the content element depicted in the predictive content generation space.

10. the user-generated shape comprises a closed shape; 9. The computer-implemented method of claim 8, wherein determining that the shape generated by the user selects the at least a portion of the content element comprises determining, by the computing system, that the closed shape generated by the user includes the at least a portion of the content element depicted within the predictive content generation space.

11. the shape generated by the user includes points corresponding to touch or click input; 9. The computer-implemented method of claim 8, wherein determining that the shape generated by the user selects the at least a portion of the content element includes determining, by the computing system, that the point generated by the user is located on the at least a portion of the content element depicted in the predictive content generation space.

12. the second tool comprises an audio brush tool, and the second machine-learned task comprises a speech recognition task; Obtaining the data indicative of the selection by the user of the predictive content element using the second tool includes: obtaining, by the computing system, data indicative of a line generated by the user within the predictive content generation space using the audio brush tool; determining, by the computing system, that the line generated by the user using the audio brush tool intersects with the predictive content element depicted within the predictive content generation space; obtaining, by the computing system, data describing a vocal utterance by the user, the vocal utterance indicating a third tool of the plurality of tools, distinct from the first tool and the second tool, the third tool being associated with a third machine learning task of the plurality of machine learning tasks; Including, Processing the data describing the predicted content elements with the machine-learned model includes: processing, by the computing system, the data describing the speech utterance with a machine-learned model trained to perform the second machine-learning task to obtain a speech recognition output that identifies the third tool; and and processing, by the computing system, the data describing the predictive content elements with a machine-learned model trained to perform the third machine learning task to obtain the second predictive content based on the speech recognition output.

13. the first machine learning task includes a content development task; 2. The computer-implemented method of claim 1, wherein the machine-learned model is trained to process the data describing the at least a portion of the content element and output content that is similar to the at least a portion of the content element.

14. the first machine learning task includes a content atomization task; the at least a portion of the content elements are associated with a concept; 10. The computer-implemented method of claim 1, wherein the machine-learned model is trained to process the data describing the at least a portion of the content element to output one or more sub-concepts of the concept.

15. the first machine learning task includes a content analysis task; 2. The computer-implemented method of claim 1, wherein the machine-learned model is trained to process the data describing the at least a portion of the content element and output a summary of the at least a portion of the content element.

16. the first machine learning task includes a prompt generation task; 2. The computer-implemented method of claim 1, wherein the machine-learned model is trained to process the data describing the at least a portion of the content element and output one or more prompts to the user related to aspects of the at least a portion of the content element.

17. Obtaining the data indicative of the selection by the user of the at least a portion of the content element comprises: selecting, by the computing system, a machine learning task from the plurality of machine learning tasks; historical user data describing the user's previous interactions within the predictive content generation space; and / or the data describing the at least a portion of the content element; selecting, based at least in part on assigning, by the computing system, the machine learning task to the first tool; 17. A computer-implemented method according to any one of claims 1 to 16, comprising:

18. The content element is: image, Video data, three-dimensional representation, Text content, a Uniform Resource Locator (URL), or Audio data, 11. The computer-implemented method of claim 10, comprising one or more of:

19. 19. The computer-implemented method of claim 1, wherein the machine-learned models include a plurality of machine-learned models that collectively process inputs in an order specified by the corresponding machine-learning tasks.

20. 20. The computer-implemented method of claim 1, wherein the method includes determining, by the computing system, the data describing the at least a portion of the content element before processing the data describing the at least a portion of the content element with the machine-learned model.

21. 20. The computer-implemented method of any preceding claim, wherein the data describing the at least a portion of the content element comprises metadata associated with the content element.

22. 22. The computer-implemented method of claim 1, wherein the predictive content generation space is a two-dimensional space.

23. 23. The computer-implemented method of claim 22, wherein the predictive content generation space is displayed on top of a separate application interface.

24. 22. The computer-implemented method of any one of claims 1 to 21, wherein the predictive content generation space is a three-dimensional augmented reality (AR) / virtual reality (VR) space.

25. 1. A computing system for generating content within a predictive content generation space via user-specified machine learning tasks, comprising: one or more processors; one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations including: obtaining data indicative of a user's selection of at least a portion of a content element depicted within the predictive content generation space using a first tool of a plurality of tools of the predictive content generation space; the plurality of tools are respectively associated with a plurality of machine learning tasks; each of the plurality of tools is operable to select at least a portion of each of one or more content elements depicted within the predictive content generation space; To get and processing data describing at least a portion of the content elements with machine-learned models to obtain predictive content, the machine-learned models being each trained to perform a first machine-learning task associated with the first tool; generating one or more predictive content elements within the predictive content generation space, the one or more predictive content elements describing the predictive content; one or more non-transitory computer-readable media, a computing system including:

26. The computing system of claim 25, wherein the operations further comprise performing a method according to any one of claims 2 to 24.

27. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations, the operations including: obtaining data indicative of a user's selection of at least a portion of a content element depicted within the predictive content generation space using a first tool of a plurality of tools of the predictive content generation space; the plurality of tools are respectively associated with a plurality of machine learning tasks; each of the plurality of tools is operable to select at least a portion of each of one or more content elements depicted within the predictive content generation space; To get and processing data describing at least a portion of the content elements with machine-learned models to obtain predictive content, the machine-learned models being each trained to perform a first machine-learning task associated with the first tool; generating one or more predictive content elements within the predictive content generation space, the one or more predictive content elements describing the predictive content; 1. One or more non-transitory computer-readable media, including:

28. 28. The one or more non-transitory computer-readable media of claim 27, wherein the operations further comprise performing a method according to any one of claims 2 to 24.