Machine learning content generation via predictive content generation space
By using user-specified machine learning tasks and multiple tools to interact with content elements in the predictive content generation space, the problem that machine learning models in the prior art are difficult to be used for tasks such as brainstorming is solved, and efficient content exploration and generation is achieved, improving user efficiency and creativity.
Patent Information
- Application Number
- CN202280100370.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-23
- Publication Date
- 2025-05-02
AI Technical Summary
It is difficult for existing technology to effectively use machine learning models for tasks such as brainstorming, content discovery, content generation, and creative exploration. Users usually use these models after identifying problems.
By using user-specified machine learning tasks in the predictive content generation space, interacting with content elements with multiple tools, selecting some or all of the content elements, and processing data describing content elements with the trained machine learning model to generate predicted content and predicted content elements.
It realizes efficient exploration and generation of content in the predictive content generation space, reduces the computing resources consumed by users, and improves user efficiency, creativity and productivity.
Smart Images

Figure CN119923671A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to content generation and prediction. More specifically, the present disclosure relates to utilizing a predictive content generation space to generate content using a machine learning model. Background Art
[0002] Advances in machine learning have led to the creation of increasingly sophisticated machine learning models. For example, large language models (LLMs) are trained on amounts of data that are much larger than the amount of data used to train conventional language models. By doing so, LLMs can be trained to perform a variety of natural language processing tasks. As another example, image processing models can be trained to perform semantic image analysis on images. In other words, in addition to identifying objects depicted in an image, these image processing models can also gain a semantic understanding of the scene itself.
[0003] Many of these models are currently used to help users perform various tasks. For example, LLMs can be used to answer questions posed by users of a search service. As another example, image processing models can be used to perform reverse image searches or provide image suggestions to users. However, in current implementations, users only utilize these models after identifying a question, so these models cannot be effectively used for brainstorming, content discovery, content generation, creative exploration, etc. Summary of the invention
[0004] Various aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned through practice of the embodiments.
[0005] An example aspect of the present disclosure relates to a computer-implemented method for performing content generation in a predictive content generation space via a user-specified machine learning task. The method includes obtaining, by a computing system including one or more computing devices, data indicating a user's selection of at least a portion of content elements depicted in the predictive content generation space using a first tool among a plurality of tools of the predictive content generation space. The plurality of tools are respectively associated with a plurality of machine learning tasks. Each of the plurality of tools is operable to select at least a portion of each of the one or more content elements depicted in the predictive content generation space. The method includes processing, by the computing system, data describing at least a portion of the content elements with a machine learning model to obtain predicted content, wherein the machine learning model is trained to perform a first machine learning task respectively associated with the first tool. The method includes generating, by the computing system, one or more predicted content elements in the predictive content generation space, wherein the one or more predicted content elements describe predicted content.
[0006] Another example aspect of the present disclosure relates to a computing system that generates content in a predictive content generation space via a machine learning task specified by a user. The computing system includes one or more processors. The computing system includes one or more non-transitory computer-readable media that store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations include obtaining data indicating a user's selection of at least a portion of a content element depicted in the predictive content generation space using a first tool among a plurality of tools in the predictive content generation space. The plurality of tools are respectively associated with a plurality of machine learning tasks. Each of the plurality of tools is operable to select at least a portion of each of the one or more content elements depicted in the predictive content generation space. The operations include processing data describing at least a portion of the content element with a machine learning model to obtain predicted content, wherein the machine learning model is trained to perform a first machine learning task associated with the first tool, respectively. The operations include one or more predicted content elements in the predictive content generation space, wherein the one or more predicted content elements describe predicted content.
[0007] Another example aspect of the present disclosure relates to one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations. The operations include obtaining data indicating a user's selection of at least a portion of content elements depicted within the predictive content generation space using a first tool among a plurality of tools in the predictive content generation space. The plurality of tools are respectively associated with a plurality of machine learning tasks. Each of the plurality of tools is operable to select at least a portion of each of one or more content elements depicted within the predictive content generation space. The operations include processing data describing at least a portion of the content elements with a machine learning model to obtain predicted content, wherein the machine learning model is trained to perform a first machine learning task respectively associated with the first tool. The operations include one or more predicted content elements within the predictive content generation space, wherein the one or more predicted content elements describe predicted content.
[0008] Other aspects of the disclosure relate to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.
[0009] These and other features, aspects and advantages of various embodiments of the present disclosure will be better understood with reference to the following description and appended claims.The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] A detailed discussion of the embodiments for those skilled in the art is set forth in this specification with reference to the accompanying drawings, in which:
[0011] Figure 1A Depicted is a block diagram of an example computing system that performs content generation within a predictive content generation space, in accordance with some implementations of the present disclosure.
[0012] Figure 1B Depicted is a block diagram of an example computing device that performs content generation within a predictive content generation space, in accordance with some implementations of the present disclosure.
[0013] Figure 1C Depicted is a block diagram of an example computing device that performs training of a machine learning model for predictive content generation, in accordance with some implementations of the present disclosure.
[0014] Figure 2A Depicted is an example layout of an interface of a predictive content generation space at a first time, according to some implementations of the present disclosure.
[0015] Figure 2B Depicted is an example interface for machine learning generation of predictive content elements within a predictive content generation space, according to some implementations of the present disclosure.
[0016] Figure 2C Depicted are example layouts of an interface of a predictive content generation space at a second time, according to some implementations of the present disclosure.
[0017] Figure 2D Depicted are content elements and predicted content elements within an interface of a predictive content generation space at a second time, according to some implementations of the present disclosure.
[0018] Figure 3A An example layout of an interface according to some other implementations of the present disclosure is shown in which an alternative user selection input is provided at a first time.
[0019] Figure 3B Depicted are predicted content elements corresponding to a selection of an entire content element within an interface at a second time, in accordance with some implementations of the present disclosure.
[0020] Figure 3C Depicted are predicted content elements corresponding to selection of the predicted content element within the interface at a third time, according to some implementations of the present disclosure.
[0021] Figure 4AAn example layout of an interface according to some other implementations of the present disclosure is shown, in which a voice brush is selected next to a spoken utterance provided by a user at a first time.
[0022] Figure 4B Depicted are predicted content elements corresponding to a selection of a content element via a voice brush tool within an interface at a second time, according to some implementations of the present disclosure.
[0023] Figure 5 Depicted are data structures that associate tools with machine learning tasks according to some implementations of the present disclosure.
[0024] Figure 6 Depicted is a block diagram of an example machine learning model according to some implementations of the present disclosure.
[0025] Figure 7 Depicted is a block diagram of an example machine learning model according to some other implementations of the present disclosure.
[0026] Figure 8 Depicted is a block diagram of an example machine learning model according to some other implementations of the present disclosure.
[0027] Fig. 9 Depicted is a flow diagram of an example method for performing content generation within a predictive content generation space according to an example implementation of the present disclosure.
[0028] Reference numerals repeated in various figures are intended to identify like features in the various implementations. DETAILED DESCRIPTION
[0029] Overview
[0030] In general, the present disclosure relates to content generation and prediction. More specifically, the present disclosure relates to using a predictive content generation space to generate content using a machine learning model. For example, a computing system (e.g., a user device, a server hosting a predictive content generation space service, etc.) can obtain data indicating that a user has selected some portion (or all) of the content elements depicted in the predictive content generation space using one of a plurality of tools. The predictive content generation space can be a two-dimensional space or a three-dimensional space (e.g., an augmented reality (AR) / virtual reality (VR) space, etc.) in which content elements are depicted. Content elements can be interface elements including images, video data, text content, uniform resource locators (URLs), audio data, etc. Users can interact with content elements via various tools. Each of these tools can correspond to different machine learning tasks. Therefore, by selecting a content element using a certain brush (e.g., by drawing a line on a content element using the brush, etc.), users can indicate the machine learning task they wish to perform.
[0031] To follow up on the previous example, a user may use a content analysis brush corresponding to a content analysis task to select a content element that includes an image of a cat. The computing system may process data describing the content element (e.g., metadata associated with the image of a cat, etc.) with a machine learning model trained to perform the content analysis task to obtain predicted content. For example, the predicted content may be textual content describing information about the breed of the cat depicted in the image, a clarifying prompt to the user corresponding to the image of the cat, etc.
[0032] The computing system may generate a predicted content element that describes the predicted content. For example, if the predicted content includes an image similar to the input image and text content about the breed of cat depicted in the input image, the computing system may generate a predicted content element that includes both the similar image and the text content within the predicted content element. In this way, the predictive content generation space may be utilized with sophisticated machine learning models to improve user efficiency, creativity, and productivity.
[0033] The implementation of the present disclosure provides a variety of technical effects and benefits. As an example technical effect and benefit, users who use conventional model implementations for content generation (e.g., search engines, reverse image search services, language processing services, etc.) must navigate between various discrete services that are only configured to perform narrowly defined tasks. For example, a user can use a reverse image search service to find an output image similar to an input image. If the user expects to obtain additional information about the entity depicted in the output image, the user will be forced to store its output image locally while trying to navigate to different services for semantic image analysis, thus unnecessarily using a large amount of computing resources (e.g., power, memory, storage, bandwidth, computing cycles, etc.). However, the implementation of the present disclosure promotes efficient exploration and generation of content by utilizing complex machine learning models in combination with a continuous predictive content generation space that provides users with various tools. By providing a more efficient content generation space, the implementation of the present disclosure can significantly reduce the amount of computing resources consumed by users.
[0034] Furthermore, implementations of the present disclosure may actually provide alternative graphical shortcuts within a graphical user interface, allowing a user to directly access and configure machine learning model tools, as well as specify inputs to the model.
[0035] Referring now to the accompanying drawings, example embodiments of the present disclosure will be discussed in greater detail.
[0036] Example Apparatus and Systems
[0037] Figure 1A A block diagram of an example computing system 100 that performs content generation within a predictive content generation space is depicted in accordance with some implementations of the present disclosure. System 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 that are communicatively coupled via a network 180.
[0038] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop computer), a mobile computing device (e.g., a smart phone or tablet computer), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0039] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or a plurality of processors operatively connected. The memory 114 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 may store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.
[0040] In some implementations, the user computing device 102 may store or include one or more machine learning models 120. For example, the machine learning model 120 may be or may otherwise include various machine learning models, such as a neural network (e.g., a deep neural network) or other types of machine learning models, including nonlinear models and / or linear models. The neural network may include a feedforward neural network, a recurrent neural network (e.g., a long short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks. Some example machine learning models may utilize attention mechanisms, such as self-attention. For example, some example machine learning models may include a multi-head self-attention model (e.g., a Transformer model). Reference Figures 5 to 7 Discuss example machine learning model 120.
[0041] In some implementations, one or more overall models 120 may be received from the server computing system 130 via the network 180, stored in the user computing device memory 114, and then used or otherwise implemented by the one or more processors 112. In some implementations, the user computing device 102 may implement multiple parallel instances of a single machine learning model 120 (e.g., to perform parallel machine learning tasks across multiple instances of the machine learning model 120).
[0042] More specifically, the machine learning model 120 may be one or more models trained to perform various machine learning tasks. For example, the machine learning model 120 may be or otherwise include a large language model (LLM) trained to perform various natural language processing tasks. For another example, the machine learning model 120 may include a semantic image processing model trained to perform semantic image analysis tasks (e.g., identifying depicted entities, scene determination, etc.). Additionally or alternatively, in some implementations, the machine learning model 120 may be or otherwise include a machine learning model pipeline, integration, etc., which includes multiple machine learning models configured to process inputs in a certain order (e.g., an order corresponding to the machine learning task).
[0043] For example, a content expansion task (e.g., a task of finding content similar to an input) may specify that a machine learning semantic image processing model is to process an input image to obtain a semantic description of the input image, and then specify that a machine learning content retrieval model is to process the semantic description to retrieve predicted content similar to the input. Thus, it should be broadly understood that the machine learning model 120 may be or may otherwise include any type of grouping or collection of machine learning models.
[0044] Additionally or alternatively, one or more machine learning models 140 may be included in or otherwise stored and implemented by a server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the machine learning model 140 may be implemented by the server computing system 130 as part of a web service (e.g., a predictive content generation space service). Thus, one or more models 120 may be stored and implemented at the user computing device 102, and / or one or more models 140 may be stored and implemented at the server computing system 130.
[0045] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or a touchpad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component may be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other devices by which a user can provide user input.
[0046] The server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or a plurality of processors operatively connected. The memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 may store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.
[0047] In some implementations, the server computing system 130 includes or is otherwise implemented by one or more server computing devices. Where the server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0048] As described above, the server computing system 130 may store or otherwise include one or more machine learning models 140. For example, the model 140 may be or may otherwise include a variety of machine learning models. Example machine learning models include neural networks or other multi-layer nonlinear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some example machine learning models may utilize attention mechanisms, such as self-attention. For example, some example machine learning models may include a multi-head self-attention model (e.g., a Transformer model). Reference Figures 5 to 7 Example model 140 is discussed.
[0049] User computing device 102 and / or server computing system 130 may train models 120 and / or 140 via interaction with training computing system 150 communicatively coupled via network 180. Training computing system 150 may be separate from server computing system 130 or may be part of server computing system 130.
[0050] The training computing system 150 includes one or more processors 152 and memory 154. The one or more processors 152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or a plurality of processors operatively connected. The memory 154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 154 may store data 156 and instructions 158, which are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes or is otherwise implemented by one or more server computing devices.
[0051] The training computing system 150 may include a model trainer 160 that trains the machine learning models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques such as, for example, error back propagation. For example, a loss function may be back propagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques may be used to iteratively update the parameters in multiple training iterations.
[0052] In some implementations, performing error back-propagation may include performing truncated back-propagation through time.The model trainer 160 may perform a variety of generalization techniques (eg, weight decay, dropout, etc.) to improve the generalization capabilities of the model being trained.
[0053] Specifically, model trainer 160 may train machine learning models 120 and / or 140 based on a set of training data 162. Training data 162 may include, for example, data sufficient to train a complex model such as an LLM or a semantic image processing model (e.g., language data, image data, etc.).
[0054] In some implementations, if the user has provided consent, the training examples may be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 may be trained by the training computing system 150 based on user-specific data received from the user computing device 102. In some cases, this process may be referred to as personalizing the model.
[0055] The model trainer 160 includes computer logic for providing the desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software that controls a general purpose processor. For example, in some implementations, the model trainer 160 includes a program file stored on a storage device, loaded into a memory, and executed by one or more processors. In other implementations, the model trainer 160 includes one or more computer executable instruction sets stored in a tangible computer readable storage medium such as RAM, a hard disk, or an optical or magnetic medium.
[0056] The network 180 may be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and may include any number of wired or wireless links. In general, communications over the network 180 may be conducted over any type of wired and / or wireless connection using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).
[0057] The machine learning models described in this specification can be used for a variety of tasks, applications, and / or use cases.
[0058] In some implementations, the input of the machine learning model of the present disclosure may be image data. The machine learning model may process the image data to generate an output. As an example, the machine learning model may process the image data to generate an image recognition output (e.g., recognition of image data, potential embedding of image data, encoded representation of image data, hash of image data, etc.). As another example, the machine learning model may process the image data to generate an image segmentation output. As another example, the machine learning model may process the image data to generate an image classification output. As another example, the machine learning model may process the image data to generate an image data modification output (e.g., a change in image data, etc.). As another example, the machine learning model may process the image data to generate an encoded image data output (e.g., an encoded representation and / or a compressed representation of image data, etc.). As another example, the machine learning model may process the image data to generate an amplified image data output. As another example, the machine learning model may process the image data to generate a prediction output.
[0059] In some implementations, the input of the machine learning model of the present disclosure may be text or natural language data. The machine learning model may process the text or natural language data to generate an output. As an example, the machine learning model may process the natural language data to generate a language encoding output. As another example, the machine learning model may process the text or natural language data to generate a potential text embedding output. As another example, the machine learning model may process the text or natural language data to generate a translation output. As another example, the machine learning model may process the text or natural language data to generate a classification output. As another example, the machine learning model may process the text or natural language data to generate a text segmentation output. As another example, the machine learning model may process the text or natural language data to generate a semantic intent output. As another example, the machine learning model may process the text or natural language data to generate an amplified text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language, etc.). As another example, the machine learning model may process the text or natural language data to generate a prediction output.
[0060] In some implementations, the input of the machine learning model of the present disclosure may be speech data. The machine learning model may process speech data to generate an output. As an example, a machine learning model may process speech data to generate a speech recognition output. As another example, a machine learning model may process speech data to generate a speech conversion output. As another example, a machine learning model may process speech data to generate a potential embedding output. As another example, a machine learning model may process speech data to generate an encoded speech output (e.g., an encoded representation and / or a compressed representation of speech data, etc.). As another example, a machine learning model may process speech data to generate an amplified speech output (e.g., speech data with a higher quality than the input speech data, etc.). As another example, a machine learning model may process speech data to generate a text representation output (e.g., a text representation of the input speech data, etc.). As another example, a machine learning model may process speech data to generate a predicted output.
[0061] In some implementations, the input of the machine learning model of the present disclosure may be latent coded data (e.g., a latent space representation of an input, etc.). The machine learning model may process the latent coded data to generate an output. As an example, the machine learning model may process the latent coded data to generate a recognition output. As another example, the machine learning model may process the latent coded data to generate a reconstruction output. As another example, the machine learning model may process the latent coded data to generate a search output. As another example, the machine learning model may process the latent coded data to generate a re-clustering output. As another example, the machine learning model may process the latent coded data to generate a prediction output.
[0062] In some implementations, the input to the machine learning model of the present disclosure may be statistical data. Statistical data may be, represent, or otherwise include data calculated and / or computed from some other data source. The machine learning model may process statistical data to generate an output. As an example, the machine learning model may process statistical data to generate an identification output. As another example, the machine learning model may process statistical data to generate a prediction output. As another example, the machine learning model may process statistical data to generate a classification output. As another example, the machine learning model may process statistical data to generate a segmentation output. As another example, the machine learning model may process statistical data to generate a visualization output. As another example, the machine learning model may process statistical data to generate a diagnostic output.
[0063] In some implementations, the input of the machine learning model of the present disclosure may be sensor data. The machine learning model may process the sensor data to generate an output. As an example, the machine learning model may process the sensor data to generate a recognition output. As another example, the machine learning model may process the sensor data to generate a prediction output. As another example, the machine learning model may process the sensor data to generate a classification output. As another example, the machine learning model may process the sensor data to generate a segmentation output. As another example, the machine learning model may process the sensor data to generate a visualization output. As another example, the machine learning model may process the sensor data to generate a diagnostic output. As another example, the machine learning model may process the sensor data to generate a detection output.
[0064] In some cases, the machine learning model can be configured to perform a task that includes encoding input data for reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task can be an audio compression task. The input can include audio data, and the output can include compressed audio data. In another example, the input includes visual data (e.g., one or more images or videos), the output includes compressed visual data, and the task is a visual data compression task. In another example, the task can include generating an embedding for the input data (e.g., input audio or visual data).
[0065] In some cases, the input includes visual data, and the task is a computer vision task. In some cases, the input includes pixel data of one or more images, and the task is an image processing task. For example, the image processing task may be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the possibility that the one or more images depict an object belonging to an object class. The image processing task may be object detection, where the image processing output identifies one or more regions in the one or more images, and for each region, identifies the possibility that the region depicts an object of interest. As another example, the image processing task may be image segmentation, where the image processing output defines the corresponding possibility of each category in a set of predetermined categories for each pixel in the one or more images. For example, the group of categories may be foreground and background. As another example, the group of categories may be object classes. As another example, the image processing task may be depth estimation, where the image processing output defines a corresponding depth value for each pixel in the one or more images. As another example, the image processing task may be motion estimation, where the network input includes multiple images, and the image processing output defines the motion of the scene depicted at the pixel between images in the network input for each pixel of one of the input images.
[0066] In some cases, the input includes audio data representing a spoken utterance, and the task is a speech recognition task. The output may include a text output mapped to the spoken utterance. In some cases, the task includes encrypting or decrypting the input data. In some cases, the task includes a microprocessor execution task, such as branch prediction or memory address translation.
[0067] Figure 1A An example computing system that can be used to implement the present disclosure is shown. Other computing systems may also be used. For example, in some implementations, the user computing device 102 may include a model trainer 160 and a training data set 162. In such implementations, the model 120 may be both trained and used locally at the user computing device 102. In some of such implementations, the user computing device 102 may implement the model trainer 160 to personalize the model 120 based on user-specific data.
[0068] Figure 1B Depicted is a block diagram of an example computing device 10 that performs content generation within a predictive content generation space, according to some implementations of the present disclosure. Computing device 10 may be a user computing device or a server computing device.
[0069] Computing device 10 includes multiple applications (e.g., application 1 to application N). Each application contains its own machine learning library and machine learning model. For example, each application can include a machine learning model. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc.
[0070] like Figure 1B As shown, each application can communicate with multiple other components of the computing device (e.g., such as one or more sensors, a context manager, a device state component, and / or additional components). In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to the application.
[0071] Figure 1C A block diagram of an example computing device 50 that performs training of a machine learning model for predictive content generation according to some implementations of the present disclosure is depicted. The computing device 50 may be a user computing device or a server computing device.
[0072] The computing device 50 includes a plurality of applications (e.g., Application 1 to Application N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a public API across all applications).
[0073] The central intelligence layer includes multiple machine learning models. For example, Figure 1C As shown, a corresponding machine learning model may be provided for each application, and the corresponding machine learning model may be managed by the central intelligence layer. In other implementations, two or more applications may share a single machine learning model. For example, in some implementations, the central intelligence layer may provide a single model for all applications. In some implementations, the central intelligence layer is included in the operating system of the computing device 50 or is otherwise implemented by the operating system.
[0074] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized data repository for the computing device 50. Figure 1C As shown, the central device data layer can communicate with multiple other components of the computing device (e.g., such as one or more sensors, context managers, device state components, and / or additional components). In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0075] Figure 2A An example layout of an interface 200 of a predictive content generation space at a first time according to some implementations of the present disclosure is depicted. Specifically, as depicted, the interface 200 of the predictive content generation space is a two-dimensional interface 200, which includes a background 201 and a toolbar 204, which includes tools 204A, 204B, 204C, 204D, and 204E. The background of the interface 200 is a solid color background, such as a whiteboard background. It should be noted that Figures 2A to 4B Only a two-dimensional whiteboard is depicted to more clearly illustrate various aspects of the present disclosure. However, implementations of the present disclosure are not limited to backgrounds or interfaces in any manner. For example, the background 202 of the interface 200 may alternatively depict a traditional blackboard background. For another example, the background 202 of the interface 200 may depict a background selected by the user.
[0076] Additionally or alternatively, in some implementations, the interface 200 can be overlaid on top of the interface of a separate application. For example, the predictive content generation space can be an operating system (OS) level feature that can be executed simultaneously with other applications. A user can browse Internet content via a web browser and then the predictive content generation space can be executed. The background 202 of the interface 200 can be the web page that the user is browsing.
[0077] The interface 200 may include a toolbar 204 that provides access to tools 204A-204E. Tools 204A-204E are brush tools that a user may utilize to generate (i.e., draw) lines, shapes, points, etc. within the interface 200 of the predictive content generation space. Figure 5 As depicted, each of the tools 204A-204E can be associated with a machine learning task in a plurality of machine learning tasks, respectively.
[0078] Go to Figure 5 , Figure 5 A data structure 500 is depicted that associates tools 204A-204E with machine learning tasks according to some implementations of the present disclosure. For example, tool 204A can be associated with content expansion task 502A. Tool 204B can be associated with content atomization task 502B. Tool 204C can be associated with content analysis task 502C. Tool 204D can be associated with prompt generation task 502D. Tool 204E can be associated with content synthesis task 502E. In some implementations, machine learning tasks 502 can each be associated with model processing instructions 504. Model processing instructions 504 can indicate a series of processing steps required to complete the corresponding machine learning task 502. For example, for an input image with content expansion task 502A, the corresponding model processing instructions 504 can indicate that the image should be processed using a semantic processing model of machine learning to obtain an intermediate output, and then the intermediate output should be processed using a retrieval model of machine learning to obtain predicted content.
[0079] In some implementations, the tools 204A-204E can be associated with user history data 506. The user history data 506 can describe the previous use of the tools 204A-204E by a particular user. For example, the user history data 506 can indicate that the user rarely uses tool 204D, but often uses tool 204A. Therefore, if a tool 204A-204E is automatically selected for the user, the computing system (e.g., server computing system 130, user computing device 102, etc.) can determine which tool of the tools 204A-204E to select based on the user history data 506.
[0080] return Figure 2A It should be noted that although tools 204A-204E are depicted as brush tools, tools 204A-204E are not necessarily limited to brush tool implementations. Instead, toolbar 204 may include any manner of tools operable to select content elements 206 within interface 200 of predictive content generation space. For example, tool 204A may instead be a box tool that a user may use to surround an item of content of interest. For another example, tool 204A may instead be a writing tool that allows a user to write instructions directly to interface 200 of predictive content generation space.
[0081] The interface 200 of the predictive content generation space may include one or more content elements 206. As depicted, the content element 206 may include a plurality of parts 206A, 206B, 206C, and 206D. Each of these parts may depict one or more entities within the content element. As previously described, the content element 206 may include any type or combination of multimedia content (e.g., audio data, text content, URLs, video data, images, three-dimensional representations, video games, web pages, summaries, live broadcasts, etc.).
[0082] At a first time T1, a user may utilize a tool of toolbar 204 to select at least a portion of content element 206 depicted within interface 200 of predictive content generation space. As depicted, the user has provided selection input 208 using tool 204B that intersects portion 206D of content element 206. Specifically, selection input 208 is a line that starts at position 208A, intersects portion 206D, and ends at position 208B. In some implementations, by intersecting portion 206D with selection input 208, the user may select portion 206D of content element 206 while excluding portions 206A-206C. Alternatively, in some implementations, by intersecting portion 206D with selection input 208, the entire content element 206 (e.g., all portions 206A-206D, etc.) may be selected.
[0083] It should be noted that although the predictive content generation space is depicted as a two-dimensional space, it is not limited to a two-dimensional space. For example, the predictive content generation space can be a virtual three-dimensional space, and the interface 200 can be displayed in a display device of an augmented reality (AR) / virtual reality (VR) device. In some implementations, the toolbar 204 may include tools for a three-dimensional environment. For example, the tool 204A may be a wand that allows the user to draw in three dimensions. For another example, the tool 204A may simulate the interaction between the user's appendage and an augmented reality element (e.g., a three-dimensional rendering, etc.) to allow the user to interact directly with the content element as an augmented reality element with their hands. For example, the content element may be a three-dimensional box indicating a video on a video sharing website. The user may provide a selection input by touching the content element. In response, the video may be played in the box, or alternatively, may be played directly to the user via a different interface (e.g., a two-dimensional interface 200, etc.). In this way, the interface 200 may switch between a two-dimensional interface and a three-dimensional interface to facilitate the user to interact with the content element using the tools of the toolbar 204.
[0084] It should be noted that, although not depicted, a user may provide content element 206 directly to interface 200 of predictive content generation space. For example, content element 206 may be an image, and a user may "drag and drop" content element 206 from a file storage system to interface 200 of predictive content generation space to initiate uploading of the image to interface 200. For another example, a user may enter text content directly into interface 200 in a free-form manner (e.g., clicking a location on interface 200 and typing may create a "text box" in which a content element including text content may be created). For another example, a user may copy a URL of a video hosted on a hosting website and paste the URL into interface 200. The video may then be displayed directly within content element 206. Thus, it should be broadly understood that interface 200 can be configured such that a user can “place,” upload, or otherwise provide content of any form or manner (e.g., videos, images, search queries, multimodal search queries, video games, AR / VR objects, URLs, web applications, virtual computing instances, etc.) to interface 200, which can then be displayed within a content element of interface 200 (e.g., content element 206).
[0085] For example, in some implementations, the content element 206 can be a query (e.g., a query image, a text query, a spoken utterance including a query, etc.). For example, the interface 200 of the predictive content generation space can allow a user to generate text content directly within the predictive content generation space. The user can then select a content element (e.g., a query) to perform a search.
[0086] Additionally or alternatively, in some implementations, content element 206 may include multiple content elements that together form a multimodal search query. For example, a user may enter a text query for "blue shoes" in a text box content element depicted in interface 200. The user may then "drag and drop" an image of a white shoe within interface 200 to form a content element that includes an image of a white shoe. The user may select both the text box content element and the image content element to provide a multimodal search query that may be processed with a machine learning model.
[0087] Figure 2B An example interface 200 for machine learning generation of predictive content elements within a predictive content generation space according to some implementations of the present disclosure is depicted. Specifically, a machine learning model 209 (e.g., Figure 1A The machine learning model 209 may be a machine learning model 120 or 140, etc., to process a portion 206D of the content element 206 selected by the user using the selection input 208 to obtain predicted content 210. The machine learning model 209 may be a machine learning model trained to perform a machine learning task associated with the tool 204B. For example, the data describing the portion 206D may be image data. The tool 204B may be associated with a machine learning task to perform content expansion (e.g., retrieving content similar to the input content, etc.). The machine learning model 209 may be or otherwise include a semantic image processing model that is trained to process an input image and retrieve (or generate) semantically similar images. The predicted content 210 may include the retrieved / generated image, or may alternatively include data indicating the image (e.g., a hyperlink to the hosting location of the image, a pointer to the storage location of the image in memory, etc.).
[0088] It should be noted that the machine learning task for which the machine learning model 209 is trained can be a task in which the output includes multiple types of media. To follow the previous example, the machine learning model 209 can include another model (or the same model) that is configured to retrieve information about the data of the image contained in the description portion 206D. For example, the semantic image processing model of the previous example can process the data of the image describing portion 206D to obtain a semantic description output. One model can use the semantic description output to retrieve similar images (or can retrieve the image in a conventional manner), while another model can generate or retrieve information about the semantic description output. For example, the portion 206D can depict ducks swimming in a pond. The semantic description output may indicate that the image depicts ducks swimming in a pond. Another model of the machine learning model 209 can generate information about the history of the duck (e.g., a large language model, etc.).
[0089] Based on the predicted content 210, the predicted content element generator 211 can generate predicted content elements 212A-212D. The predicted content elements 212A-212D content elements describe the predicted content 210. For example, the predicted content 210 may include an image. The predicted content element 212A may depict or otherwise include an image. For another example, the predicted content 210 may include textual content. The predicted content element 212A may include textual content, a summary of the textual content, or a link to a location hosting the textual content. For another example, the predicted content 210 may include a cloud-based video game. The predicted content element 210A may include a cloud-based video game configured to execute when the user selects the predicted content element 212A.
[0090] It should be noted that the operations of the machine learning model 209 and the predictive content element generator 211 are depicted as occurring outside of the interface 200 to merely indicate that these operations are not depicted within the interface 200. Therefore, the depicted locations of the machine learning model 209 and the predictive content element generator 211 should not be interpreted as indicating which computing device(s) are used to perform the operations of the machine learning model 209 and the predictive content element generator 211.
[0091] Figure 2C An example layout of an interface 200 of a predictive content generation space at a second time according to some implementations of the present disclosure is depicted. Specifically, at time T2, predictive content elements 212A, 212B, 212C, and 212D are generated and depicted within the interface 200 of the predictive content generation space. In some implementations, a connection interface element 212 may be generated within the predictive content generation space. The connection interface element 212 may depict a connection between the predictive content elements 212A-212D and the portion 206D of the content element 206.
[0092] It should be noted that in some implementations, the predicted content elements 212A-212D can be generated and depicted at the location within the interface 200 where the user has completed the selection input 208. For example, Figure 2A As depicted, the user has concluded their selection input 208 at location 208B. Thus, predicted content elements 212A-212D are generated and depicted at substantially the same location as location 208B. Alternatively, in some implementations, the locations at which predicted content elements 212A-212D are generated and depicted can be determined in some other manner (e.g., based on user preferences, user history data, content types of predicted content elements 212A-212D, etc.).
[0093] At time T2, portion 206D of content element 206 selected by user selection input 208 has been processed with the machine learning model. More specifically, data describing portion 206D has been processed with the machine learning model to obtain predicted content. The machine learning model can be a model trained to perform a machine learning task associated with tool 204B. Predicted content elements 210A, 210B, 210C, and 210D can be generated based on the predicted content.
[0094] Figure 2D Depicts content elements and predicted content elements within an interface 200 of a predictive content generation space at a second time according to some implementations of the present disclosure. Specifically, it should be noted that Figure 2C Only presented Figure 2C , which shows the content depicted within content elements 206 and 212A-212D, rather than the layout of content elements 206 and 212A-212D. For example, content element 206 includes an image depicting a top-down layout of a room. The room includes a sofa, a television, a table, a plant, and a ping-pong table.
[0095] As depicted, the plant is located within portion 206D of content element 206 selected by the user using tool 204B by selecting input 208. Thus, predicted content elements 212A-212D include content corresponding to a machine learning task associated with tool 204B. In this example, tool 204B may be a content expansion tool that is a task to find content similar to the content of portion 206D. Following this example, the plant depicted in portion 206D may be a sunflower. Predicted content 210 retrieved or generated using machine learning model 209 may be related to sunflowers. For example, content element 212A may include information about a sunflower plant (e.g., retrieved from an online dictionary, synthesized using a large language model, etc.). Predicted content element 212B may include images related to caring for sunflowers and links to videos hosted on video sharing websites. Predicted content element 212C may be an image of a sunflower. Predicted content element 212D may be a grouping of concept tags related to decorating a room (i.e., the purpose of sunflowers). For example, if the user selects the furniture content tag from the predicted content element 212D (e.g., via touch input, click input, etc.), a second content element may be generated that includes predicted content about furniture that may be related to furniture depicted in other portions of the content element 206.
[0096] in addition, Figure 2DData 214 describing content element 206 is shown. Specifically, the data describing content element 206 is metadata 214 about images that collectively depict an overhead view of a room. For example, metadata 214 describes each entity depicted within the room (e.g., a television, a sofa, a table, a plant, etc.). Additionally, metadata 214 may describe a semantic view of the image (e.g., the image depicts an overhead view of a family room). In some implementations, metadata may already be included in content element 206. For example, a content element may include an image depicting a room, and the image may include metadata 214. Alternatively, in some implementations, metadata 214 may be determined. For example, content element 206 may be processed using machine learning model 209 to determine metadata 214.
[0097] It should be noted that while metadata 214 describes the entire content element 206, it may alternatively describe only relevant portions of content element 206. To follow the previous example, after a user selects portion 206D using selection input 208, metadata 214 may be determined for portion 206D.
[0098] Figure 3A 2 shows an example layout of an interface 200 according to some other implementations of the present disclosure, in which an alternative user selection input is provided at the first time. Figure 2A As depicted, selection input 208 from the user may be a line that intersects a portion of content element 206. However, the user is not limited to such selection inputs, nor is the user limited to selecting portions of content element 206. For example, as depicted, the user may provide selection input 302 as a closed shape input. Specifically, the user may use brush tool 204B to draw a closed shape around content element 206 to perform selection input 302 that selects the entire content element 206. In this manner, the user may indicate interest in all portions of content element 206 rather than a specific portion.
[0099] Alternatively, in some implementations, the user can perform a click input 303. Specifically, the user can select the brush tool 204B and then draw a "point" (i.e., click or touch a location on the interface 200) to perform a selection input 302 that selects a portion or all of the content element 206. For example, the user can click the center of the content element 206 to indicate interest in the entire content element 206. For another example, the user can click to provide a selection input 303 farther from the center of the content element 206 to indicate interest in a specific portion of the content element 206. In some implementations, the user's intent associated with the selection input 303 can be determined based on user history data and / or the content of the content element 206.
[0100] Figure 3BDepicted are predicted content elements corresponding to a selection of entire content element 206 at a second time within interface 200, according to some implementations of the present disclosure. Specifically, at time T2, following entire content element 206, predicted content elements 304A, 304B, 304C, and 304D are generated and depicted within interface 200, as described with respect to predicted content elements 212A-212D.
[0101] For example, predictive content element 304A may include a summary of a website indicating whether family-related programming is available for streaming. Predictive content element 304B may include a picture of a table tennis racket that is popular among professional players. Predictive content element 304C may include a picture of a table tennis racket that is popular among professional players. Figure 2D Predicted content element 304D may include the same sunflower image as predicted content element 212C. Figure 2D The predicted content element 212D is a grouping of the same concept tags related to decorating a room. Figure 3B Also depicted is a second selection input 306 made by the user using tool 204B. Additionally, the user has provided the second selection input 306 using brush tool 204B. The second selection input 306 selects the predictive content element 304B, which depicts a table tennis racket that is popular among professional players.
[0102] Figure 3C Depicted is a predicted content element corresponding to a selection of predicted content element 304B in interface 200 at a third time according to some implementations of the present disclosure. Specifically, at time T3, second content elements 308A, 308B, and 308C may be generated and depicted in interface 200, as described with respect to Figure 2B As described. Since the selection input 306 has selected the content element 304B depicting an image of a table tennis racket using the content expansion tool 204B, the second content elements 308A-308C can be related to table tennis. For example, the second predicted content element 308A can include a prompt that, if selected by the user, can perform a search for table tennis coaches local to the user. The second predicted content element 308B is a video on a video hosting website related to a table tennis tournament. The second predicted content element 308C is an article related to table tennis equipment.
[0103] In some implementations, additional connection elements 212 can be generated to link the second predicted content elements 308A-308C to the content element 304B. In this way, the predicted content generation space can be used as a continuous surface for users to creatively explore content and ideas within a single space.
[0104] Figure 4AAn example layout of an interface 200 according to some other implementations of the present disclosure is shown in which a voice brush is selected next to a spoken utterance provided by a user at a first time. Specifically, the user can alternatively select tool 204C instead of selecting tool 204B. Tool 204C can correspond to a voice brush tool. The voice brush tool can be configured to select content elements or portions of content elements in the same manner as brushes 204A-B and 204D-E. However, the machine learning task associated with the voice brush tool 204C can be specified by a spoken utterance 404 provided by the user.
[0105] For example, a user may use the voice brush tool 204C to provide a selection input 402 that selects the entire content element 206. At the same time, the user may also provide a spoken utterance 404 indicating an interest in what type of television is depicted within the content element 206. In this manner, the user may indicate a particular portion or entity of interest within the content element 206 via the spoken utterance 404. Furthermore, the user may indicate which machine learning tasks should be performed via the spoken utterance 404. For example, by asking “What kind of television is this?” the user may indicate a desired content analysis task.
[0106] In some implementations, the machine learning task indicated by the spoken utterance 404 may be a machine learning task associated with another tool of the toolbar 204. To follow the previous example, the content analysis machine learning task may be associated with the tool 204E. However, by indicating the content analysis task via the spoken utterance 404, the content analysis task may be performed without the user manually selecting the tool 204E.
[0107] In some implementations, the machine learning task indicated by spoken utterance 404 may not be an explicit match for an existing machine learning task. For example, a user may simply indicate an interest in TV without a corresponding machine learning task (e.g., "that's a nice TV", etc.). In response, a large language model of machine learning model 209 may be utilized to determine an expected or optimal machine learning task.
[0108] In some implementations, the machine learning task can be selected based on user history data describing the user's previous interactions within the predictive content generation space. For example, the user history data can indicate that in a previous interaction, the user has used the content expansion tool 204B to provide a selection input similar to the selection input 402. Therefore, the content expansion task can be selected to be assigned to the voice brush tool 204C.
[0109] Additionally or alternatively, in some implementations, a machine learning task may be selected based on data describing at least a portion of a content element. Figure 2DThe metadata 214 of the content element 206 may indicate that the content element 206 includes an image depicting a television. The user history data may indicate that the user has previously selected the image of the television using the content expansion tool 204B. Alternatively, the generalized user history data may indicate that most users who select the image of the television use the content expansion tool 204B to make the selection. In this way, a machine learning task can be assigned to the voice brush tool regardless of whether the task is explicitly indicated in the spoken utterance 404.
[0110] Figure 4B Depicted are predicted content elements corresponding to the selection of content element 206 in interface 200 via a voice brush tool at a second time according to some implementations of the present disclosure. Specifically, as shown, content elements 406A, 406B, and 406C can be retrieved / generated and depicted in interface 200 based on voice brush tool 204C and spoken utterance 404. In this example, a content expansion task has been assigned to voice brush tool 204C based on spoken utterance 404. For example, predicted content element 406A may include information about a market where a television set may be purchased. Predicted content element 406B may be information retrieved from an online dictionary and summarized in a content element. Predicted content element 406C may be a summary or highlight of information retrieved from an online forum.
[0111] It should be noted that although only a single user is depicted selecting content elements in this example, implementations of the present disclosure are not limited to a single user. Instead, the predictive content generation space can be utilized by multiple users. For example, the predictive content generation space can be a collective space in which multiple users can interact independently. For example, a first user can generate a "web" of content elements at a first location in the predictive content generation space, and can then move to a second location in the predictive content generation space to generate a new "network" of content elements. Later, a second user may explore the predictive content generation space and discover the "network" of content elements of the first user at the first location. The second user can then continue to expand the content elements at the first location, or can move to a different location to create a new network of content elements.
[0112] As another example, multiple users can interact with content elements simultaneously. For example, two users can use two different brushes to select content elements within the predicted content generation space (e.g., one user "brushes" to the right so that additional content elements are generated on the right, while the other user "brushes" to the left so that additional content elements are generated on the left). In this way, multiple users can jointly create a network of content elements when exploring predicted content in a collaborative manner.
[0113] As yet another example, the state of a predictive content generation space can be saved and distributed to other users of the predictive content generation space. For example, a user can generate content within an instance of a predictive content generation space. A user can save the state of a predictive content generation space. The state of a predictive content generation space may include any content elements generated within the space, the location of the predictive content within the space, and any connection elements that exist for linking content elements within the space. The saved state of the predictive content generation space can be provided to another user who has executed a separate instance of the predictive content generation space. The other user can load the saved state of the predictive content generation space to obtain the saved content elements, locations, connection elements, etc. In this way, a session in which a user creates and explores predictive content within the predictive content generation space can be shared with other users.
[0114] Example Model Layout
[0115] Figure 6 A block diagram of an example machine learning model 600 (e.g., machine learning model 209, machine learning model 120, machine learning model 140, etc.) according to some implementations of the present disclosure is depicted. In some implementations, the machine learning model 600 is trained to receive a set of input data 604 describing at least a portion of a content element, and as a result of receiving the input data 604, provide output data 606 that is predicted content or otherwise indicates predicted content.
[0116] In some implementations, the machine learning model 600 may include a task-specific portion 602 operable to perform certain machine learning tasks. For example, the machine learning model 600 may be a large language model trained to perform multiple machine learning tasks. The task-specific portion 602 may be a portion trained to perform a single task. For example, the task-specific portion 602 may be trained to generate a summary of information to include in a content element.
[0117] Figure 7 A block diagram of an example machine learning model 700 according to some other implementations of the present disclosure is depicted. The machine learning model 700 and Figure 6 The machine learning model 600 is similar to the machine learning model 600, except that the machine learning model 700 further includes a preprocessing model 702. The preprocessing model can be any model that is trained to process the input data 604 to generate an intermediate output 704 that can be processed using the task-specific model 602.
[0118] To follow the previous example, the task-specific portion 602 can be configured to summarize the retrieved information for inclusion in a content element of the output data 606. The input data 604 can be data describing a content element including an image. The pre-processing model can be a model trained to semantically process an image to output an intermediate output 704 including a semantic description of the image. The intermediate output 704 (i.e., the semantic description of the image) can be processed with the task-specific model 602 to obtain the output data 606 (e.g., including the summarized predicted content).
[0119] Figure 8 A block diagram of an example machine learning model 800 according to some other implementations of the present disclosure is depicted. The machine learning model 800 and Figure 7 The machine learning model 700 is similar to the machine learning model 800, except that the machine learning model 800 further includes a task selection model 803. The task selection model 803 can be a model trained to process the context information 804 to output a selection of a machine learning task.
[0120] For example, context information 804 may include user history data describing users of the predictive content generation space, as described above. Additionally, in some implementations, context information 804 may describe the context in which the predictive content generation space is being used (e.g., whether it is being used in a collaborative manner, based on the theme of all content elements within the space, etc.).
[0121] In some implementations, the task selection model 803 can process the context information 804 to obtain a task selection output 806. In addition, in some implementations, the task selection model can also process input data 604 describing the selected content element. The task selection output 806 can be provided to the pre-processing model 702 and the task-specific model 602. To follow the previous example, the task selection model 803 can process the context information 804 indicating that the user prefers to use the content expansion tool. The task selection model 803 can also process the input data 604 indicating that the content element includes image data (e.g., image data can be strongly related to the content expansion task).
[0122] The task selection model 803 can provide a task selection output 803 to the pre-processing model 702 and the task-specific model 602. For example, the pre-processing model 702 can be a large model configured to be trained to perform multiple pre-processing tasks. Based on the task selection output 806, one of the multiple pre-processing tasks can be indicated to the pre-processing model 702. Similarly, the task-specific model 602 can be one of multiple task-specific models, or can be a large model trained to perform multiple tasks. Based on the task selection output 806, a model can be selected as the task-specific model 602, or a task can be indicated to the task-specific model 602. In this way, multiple conventional and large machine learning models can be utilized to perform multiple processing operations that facilitate the predictive content generation space.
[0123] Example Method
[0124] Fig. 9 A flowchart of an example method for performing content generation within a predictive content generation space is depicted according to an example implementation of the present disclosure. Fig. 9 The steps performed in a specific order are depicted for the purpose of illustration and discussion, but the method of the present disclosure is not limited to the specific illustrated order or arrangement. The various steps of method 900 may be omitted, rearranged, combined and / or adjusted in various ways without departing from the scope of the present disclosure.
[0125] At 902, a computing system obtains data indicating a selection of a content element depicted within a predictive content generation space using a first tool of a plurality of tools. Specifically, the computing system obtains data indicating a selection of at least a portion of the content elements depicted within the predictive content generation space by a user using a first tool of a plurality of tools of the predictive content generation space. The plurality of tools may be associated with a plurality of machine learning tasks, respectively. Each of the plurality of tools may be operable to select at least a portion of each of one or more content elements depicted within the predictive content generation space.
[0126] In some implementations, the first machine learning task may include a content expansion task, where the machine learning model is trained to process data describing at least a portion of the content element and output content similar to at least a portion of the content element.
[0127] In some implementations, the first machine learning task includes a content analysis task, and the machine learning model is trained to process data describing at least a portion of the content elements and output a summary of at least a portion of the content elements.
[0128] In some implementations, the first machine learning task includes a prompt generation task, and the machine learning model is trained to process data describing at least a portion of the content element and output one or more prompts related to various aspects of at least a portion of the content element to the user.
[0129] In some implementations, obtaining data indicating a user selection of at least a portion of a content element further comprises selecting a machine learning task from a plurality of machine learning tasks. The task may be selected based at least in part on historical user data describing a user's prior interactions within the predictive content generation space and / or data describing at least a portion of the content element. The computing system may assign the machine learning task to the first tool.
[0130] In some implementations, the content elements include one or more of images, video data, three-dimensional representations, textual content, URLs, audio data, video games, and the like.
[0131] In some implementations, to obtain data indicating a user selection of at least a portion of a content element, the computing system may obtain data indicating a user selection of a first portion of an image. The first portion of the image may depict a first entity, and the second portion of the image may depict a second entity different from the first entity.
[0132] In some implementations, before processing the data describing at least a portion of the content element using the machine learning model, the computing system may determine the data describing at least a portion of the content element. In some implementations, the data describing at least a portion of the content element may include metadata associated with the content element.
[0133] In some implementations, the predictive content generation space is a two-dimensional space. In some implementations, the predictive content generation space is displayed on the interface of a separate application. In some implementations, the predictive content generation space is a three-dimensional augmented reality (AR) / virtual reality (VR) space.
[0134] In some implementations, the data indicating the selection of the content element may be or otherwise include data indicating a multimodal search query. For example, the content element selected by the user may be a text query. The user may also select a content element that is an image while selecting the text query.
[0135] At 904, the computing system processes the data describing the content element with the machine learning model to obtain predicted content. Specifically, the computing system processes the data describing at least a portion of the content element with the machine learning model to obtain predicted content. The machine learning model can be trained to perform first machine learning tasks respectively associated with the first tool. In some implementations, the predicted content can include predicted content of a first content type corresponding to the first machine learning task.
[0136] In some implementations, the first machine learning task may be a multimodal search task. To follow the previous example, a user may select multiple content elements to form a multimodal search query (e.g., selecting a text query, image and video data, etc.). A machine learning model may be trained to retrieve search results based on the multimodal query, or to facilitate multimodal result retrieval (e.g., to generate embeddings that can be used to retrieve results from a multimodal search space, etc.).
[0137] In some implementations, processing data describing at least a portion of a content element includes processing data describing a first portion of the content element with a machine learning model to obtain predicted content, the predicted content including information identifying one or more images that are semantically similar to the first portion of the image of the content element.
[0138] In some implementations, the machine learning model may include multiple machine learning models that jointly process inputs in an order specified by the corresponding machine learning tasks.
[0139] In some implementations,
[0140] At 906, the computing system generates one or more predicted content elements within the predictive content generation space. The one or more predicted content elements may describe the predicted content.
[0141] In some implementations, the first machine learning task may be a machine learning semantic image retrieval task, and the content element may include an image. To process the data describing at least a portion of the content element, the computing system may process the data describing at least a portion of the content element with a machine learning model to obtain predicted content, the predicted content including information identifying one or more images that are semantically similar to the image of the content element. To generate the one or more predicted content elements, the computing system may generate one or more predicted content elements, each including one or more images, within a predictive content generation space.
[0142] In some implementations, a computing system may obtain data indicating a user's selection of a predictive content element from one or more predictive content elements depicted in a predictive content generation space using a second tool from among a plurality of tools that is different from the first tool. The computing system may process the data describing the predictive content element with a machine learning model to obtain a second predictive content. The machine learning model may be trained to perform a second machine learning task respectively associated with the second tool. The computing system may generate one or more second predictive content elements in the predictive content generation space. The one or more second predictive content elements may describe the second predictive content.
[0143] In some implementations, the computing system may generate connection elements within the predictive content generation space that depict connections between the predicted content element and one or more second predicted content elements.
[0144] In some implementations, the machine learning model includes a large language model trained to perform both the first machine learning task and the second machine learning task.
[0145] In some implementations, the first tool includes a brush tool. Obtaining data indicating a user selection of at least a portion of the content element may include obtaining data indicating a shape generated by the user within the predictive content generation space using the brush tool, and determining that the shape generated by the user selects at least a portion of the content element depicted within the predictive content generation space.
[0146] In some implementations, the shape generated by the user includes a line. Determining that the shape generated by the user selects at least a portion of the content element can include determining that the line generated by the user intersects at least a portion of the content element depicted within the predictive content generation space.
[0147] In some implementations, the shape generated by the user includes an enclosed shape.Determining that the shape generated by the user selects at least a portion of the content element can include determining that the enclosed shape generated by the user includes at least a portion of the content element depicted within the predictive content generation space.
[0148] In some implementations, the shape generated by the user includes a point corresponding to the touch input or the click input. Determining that the shape generated by the user selects at least a portion of the content element may include determining that the point generated by the user is located at at least a portion of the content element depicted within the predictive content generation space.
[0149] In some implementations, the second tool may include a voice brush tool, and the second machine learning task may include a speech recognition task. Obtaining data indicating a user's selection of a predicted content element using the second tool may include obtaining data indicating a line generated by the user using the voice brush tool within the predictive content generation space. The computing system may determine that the line generated by the user using the voice brush tool intersects the predicted content element depicted within the predictive content generation space. The computing system may obtain data describing the user's spoken utterance. The spoken utterance may indicate a third tool among multiple tools that is different from the first tool and the second tool. The third tool may be associated with a third machine learning task among multiple machine learning tasks. Processing data describing the predicted content element with a machine learning model may include processing data describing the spoken utterance with a machine learning model trained to perform the second machine learning task to obtain a speech recognition output identifying the third tool. Based on the speech recognition output, the computing system may process data describing the predicted content element with a machine learning model trained to perform the third machine learning task to obtain a second predicted content.
[0150] Additional public content
[0151] The technology discussed herein relates to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for multiple possible configurations, combinations, and partitioning of tasks and functions between and among components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system, or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0152] Although the subject matter has been described in detail with respect to various specific example embodiments of the subject matter, each example is provided by way of explanation rather than limitation of the present disclosure. Those skilled in the art may easily produce changes, modifications, and equivalents to such embodiments after understanding the foregoing. Therefore, the present disclosure does not exclude such modifications, variations, and / or additions to the subject matter that would be readily apparent to those of ordinary skill in the art. For example, a feature shown or described as part of one embodiment may be used together with another embodiment to produce yet another embodiment. Therefore, the present disclosure is intended to encompass such changes, variations, and equivalents.
Claims
1. A computer-implemented method for performing content generation in a predictive content generation space via a user-specified machine learning task, comprising: Data indicating a user's selection of at least a portion of content elements depicted within a predictive content generation space using a first tool of a plurality of tools of the predictive content generation space is obtained by a computing system including one or more computing devices, wherein: The plurality of tools are respectively associated with a plurality of machine learning tasks; and Each of the plurality of tools is operable to select at least a portion of each of one or more content elements depicted within the predictive content generation space; processing, by the computing system, the data describing the at least a portion of the content element using a machine learning model to obtain predicted content, wherein the machine learning model is trained to perform a first machine learning task respectively associated with the first tool; and One or more predicted content elements are generated by the computing system within the predictive content generation space, wherein the one or more predicted content elements describe the predicted content.
2. The computer-implemented method of claim 1, wherein: The predicted content includes predicted content of a first content type corresponding to the first machine learning task.
3. The computer-implemented method of claim 2, wherein: The first machine learning task is a machine learning semantic image retrieval task; and The content element includes an image; wherein processing the data describing the at least a portion of the content element comprises processing, by the computing system, the data describing the at least a portion of the content element with the machine learning model to obtain predicted content, the predicted content comprising information identifying one or more images that are semantically similar to the image of the content element; and The generating of the one or more predicted content elements includes generating, by the computing system, one or more predicted content elements in the predictive content generation space, each of which includes the one or more images.
4. The computer-implemented method of claim 3, wherein: Obtaining the data indicative of a selection of the at least a portion of the content element by the user comprises obtaining, by the computing system, data indicative of a selection of a first portion of the image by the user, wherein the first portion of the image depicts a first entity and a second portion of the image depicts a second entity different from the first entity; Wherein, processing the data describing at least a portion of the content element includes processing, by the computing system, the data describing the first portion of the content element using the machine learning model to obtain predicted content, wherein the predicted content includes information identifying one or more images that are semantically similar to the first portion of the image of the content element.
5. The computer-implemented method of any one of claims 1 to 4, wherein: The method further comprises: Data is obtained, by the computing system, indicating a selection by the user of a predictive content element of the one or more predictive content elements depicted within the predictive content generation space using a second tool of the plurality of tools that is different from the first tool. processing, by the computing system, the data describing the predicted content element using a machine learning model to obtain second predicted content, wherein the machine learning model is trained to perform second machine learning tasks respectively associated with the second tools; and One or more second predicted content elements are generated by the computing system within the predictive content generation space, wherein the one or more second predicted content elements describe the second predicted content.
6. The computer-implemented method of claim 5, wherein: The method further includes generating, by the computing system, a connection element within the predictive content generation space, the connection element depicting a connection between the predicted content element and the one or more second predicted content elements.
7. The computer-implemented method of any one of claims 5-6, wherein: The machine learning model includes a large language model trained to perform both the first machine learning task and the second machine learning task.
8. The computer-implemented method of any one of claims 1 to 7, wherein: The first tool includes a brush tool; and Wherein, obtaining the data indicating the selection of the at least one portion of the content elements by the user comprises: obtaining, by the computing system, data indicating a shape generated by the user using the brush tool within the predictive content generation space; and Determining, by the computing system, that the shape generated by the user selects the at least a portion of the content element depicted within the predictive content generation space.
9. The computer-implemented method of claim 8, wherein: The shape generated by the user includes a line, and wherein determining that the shape generated by the user selects at least a portion of the content element includes determining, by the computing system, that the line generated by the user intersects the at least a portion of the content element depicted within the predictive content generation space.
10. The computer-implemented method of claim 8, wherein: The shape generated by the user includes a closed shape, and wherein determining that the shape generated by the user selects at least a portion of the content element includes determining, by the computing system, that the closed shape generated by the user includes at least a portion of the content element depicted within the predictive content generation space.
11. The computer-implemented method of claim 8, wherein: The shape generated by the user includes a point corresponding to a touch input or a click input, and wherein determining that the shape generated by the user selects at least a portion of the content element includes determining, by the computing system, that the point generated by the user is located at at least a portion of the content element depicted within the predictive content generation space.
12. The computer-implemented method of any one of claims 4 to 8, wherein: The second tool includes a speech brush tool, and the second machine learning task includes a speech recognition task; Obtaining the data indicative of the selection of the predicted content element by the user using the second tool comprises: obtaining, by the computing system, data indicating a line generated by the user within the predictive content generation space using the voice brush tool; determining, by the computing system, that the line generated by the user using the voice brush tool intersects the predicted content element depicted within the predictive content generation space; and obtaining, by the computing system, data describing a spoken utterance of the user, wherein the spoken utterance indicates a third tool of the plurality of tools that is different from the first tool and the second tool, wherein the third tool is associated with a third machine learning task of the plurality of machine learning tasks; and Wherein, processing the data describing the predicted content element with the machine learning model comprises: processing, by the computing system, the data describing the spoken utterance with a machine learning model trained to perform the second machine learning task to obtain a speech recognition output identifying the third tool; and Based on the speech recognition output, the computing system processes the data describing the predicted content element using a machine learning model trained to perform the third machine learning task to obtain the second predicted content.
13. The computer-implemented method of claim 1, wherein: The first machine learning task includes a content expansion task; and The machine learning model is trained to process the data describing the at least a portion of the content element and output content similar to the at least a portion of the content element.
14. The computer-implemented method of claim 1, wherein: The first machine learning task includes a content atomization task; The at least a portion of the content elements are associated with a concept; and The machine learning model is trained to process the data describing the at least a portion of the content element and output one or more sub-concepts of the concept.
15. The computer-implemented method of claim 1, wherein: The first machine learning task includes a content analysis task; and The machine learning model is trained to process the data describing the at least a portion of the content element and output a summary of the at least a portion of the content element.
16. The computer-implemented method of claim 1, wherein: The first machine learning task comprises a prompt generation task; and The machine learning model is trained to process the data describing the at least a portion of the content element and output one or more prompts to the user related to various aspects of the at least a portion of the content element.
17. The computer-implemented method of any one of claims 1 to 16, wherein: Obtaining the data indicating the selection of the at least a portion of the content elements by the user further comprises: The computing system selects a machine learning task from the plurality of machine learning tasks based at least in part on: Historical user data describing prior interactions of the user within the predictive content generation space; and / or the data describing the at least a portion of the content element; and assigning, by the computing system, the machine learning task to the first tool.
18. The computer-implemented method of claim 10, wherein: The content elements include one or more of the following: image; Video data; Three-dimensional representation; Text content; Uniform Resource Locator (URL); or Audio data.
19. The computer-implemented method of any one of claims 1 to 18, wherein: The machine learning model includes multiple machine learning models, which jointly process inputs in an order specified by the corresponding machine learning tasks.
20. The computer-implemented method of any one of claims 1-19, wherein: Prior to processing the data describing the at least a portion of the content element with the machine learning model, the method includes determining, by the computing system, the data describing the at least a portion of the content element.
21. The computer-implemented method of any one of claims 1-19, wherein: The data describing the at least a portion of the content element comprises metadata associated with the content element.
22. The computer-implemented method of any one of claims 1 to 21, wherein: The predictive content generation space is a two-dimensional space.
23. The computer-implemented method of claim 22, wherein: The predictive content generation space is displayed on an interface of a separate application.
24. The computer-implemented method of any one of claims 1-21, wherein: The predictive content generation space is a three-dimensional augmented reality AR / virtual reality VR space.
25. A computing system for performing content generation in a predictive content generation space via a user-specified machine learning task, comprising: one or more processors; as well as one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the computing system to perform operations comprising: obtaining data indicative of a user selection of at least a portion of content elements depicted within a predictive content generation space using a first tool of a plurality of tools of the predictive content generation space, wherein: The plurality of tools are respectively associated with a plurality of machine learning tasks; and Each of the plurality of tools is operable to select at least a portion of each of one or more content elements depicted within the predictive content generation space; processing the data describing the at least a portion of the content element with a machine learning model to obtain predicted content, wherein the machine learning model is trained to perform a first machine learning task respectively associated with the first tool; and One or more predicted content elements are generated within the predictive content generation space, wherein the one or more predicted content elements describe the predicted content.
26. The computing system of claim 25, wherein: The operations further include performing the method of any one of claims 2-24.
27. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations comprising: obtaining data indicative of a user selection of at least a portion of content elements depicted within a predictive content generation space using a first tool of a plurality of tools of the predictive content generation space, wherein: The plurality of tools are respectively associated with a plurality of machine learning tasks; and Each of the plurality of tools is operable to select at least a portion of each of one or more content elements depicted within the predictive content generation space; processing the data describing the at least a portion of the content element with a machine learning model to obtain predicted content, wherein the machine learning model is trained to perform a first machine learning task respectively associated with the first tool; and One or more predicted content elements are generated within the predictive content generation space, wherein the one or more predicted content elements describe the predicted content.
28. The one or more non-transitory computer-readable media of claim 27, wherein: The operations further include performing the method of any one of claims 2-24.