Content management tool for capturing and generatively transforming content items

By using content management tools and machine learning models to automatically or user-definedly transform content items, the challenges of translation, correction, and adaptation during the pasting process are solved, improving the convenience and efficiency of the pasting process.

CN122095347APending Publication Date: 2026-05-26MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MICROSOFT TECHNOLOGY LICENSING LLC
Filing Date
2024-10-30
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

When users paste content from one location to another, they find it difficult to effectively translate, correct, adapt, and modify it, making the pasting process inconvenient.

Method used

By employing content management tools and utilizing machine learning models such as generative large language models (LLM), transformer models, diffusion models, or multimodal models, content items can be automatically or user-defined and transformed, presented and edited in different user interface elements, and finally pasted into the requested location.

Benefits of technology

It enables automatic or manual translation, correction, and adaptation of content items during the pasting process, improving the convenience and efficiency of the pasting process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122095347A_ABST
    Figure CN122095347A_ABST
Patent Text Reader

Abstract

Systems and methods for transforming captured content items are provided. Specifically, a computing device can receive a capture request for capturing content items; in response to the capture request, capture the content items and provide them in a first user interface element of a content management tool; apply a generative transformation function to the content items to generate transformed content items; write the transformed content items into a second user interface element of the content management tool; receive a paste request for pasting the transformed content items at a requested location; and in response to the paste request, provide the transformed content items at the requested location.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Computing devices encompass a variety of productivity tools and information that facilitate the completion of various tasks, including copying and pasting content items between different devices and applications. For example, clipboard tools allow users to copy and store content items (e.g., images and text) from their original location and paste the copied content items to a new location. However, it can be challenging for users to conveniently and efficiently transform (e.g., translate, correct, adapt, and / or modify) the copied content items before pasting them to a new location.

[0002] The aspects disclosed herein have been made in relation to these and other general considerations. Furthermore, although relatively specific problems may have been discussed, it should be understood that the examples are not limited to solving specific problems identified in the background of this disclosure or elsewhere. Summary of the Invention

[0003] According to examples in this disclosure, a content management tool allows a user to capture and generatively transform content items, and then copy and paste the transformed content items to a new location. When a user captures a content item, the content management tool transforms the content item by applying generative transformation features (e.g., translation, correction, adaptation, and / or revision) to transform the content item using a generative large language model (LLM), a transformer model, a diffusion model or a multimodal model, other types of machine learning models, or a combination of models. For example, a generative transformation feature is a natural language prompt describing one or more tasks to be performed on the content item to generate the transformed content item. A generative transformation feature can be automatically selected based on a previously selected generative transformation feature. Alternatively, a generative transformation feature can be selected from a list of predefined generative transformation features or defined by the user.

[0004] According to at least one example of this disclosure, a method for transforming captured content items is provided. The method may include: receiving a capture request for capturing the content item; upon receiving the capture request, capturing the content item and providing the content item in a first user interface element of a content management tool; applying a generative transformation function to the content item to generate a transformed content item; writing the transformed content item into a second user interface element of the content management tool; receiving a paste request for pasting the transformed content item at a requested location; and in response to receiving the paste request, providing the transformed content item at the requested location.

[0005] According to at least one example of this disclosure, a computing device for transforming captured content items is provided. The computing device may include: a processor; and a memory storing a plurality of instructions, which, when executed by the processor, cause the computing device to: receive a capture request for capturing content items; in response to the capture request, capture the content items and provide the content items in a first user interface element of a content management tool; apply a generative transformation function to the content items to generate transformed content items; write the transformed content items into a second user interface element of the content management tool; receive a paste request for pasting the transformed content items at a requested location; and in response to the paste request, provide the transformed content items at the requested location.

[0006] According to at least one example of this disclosure, a method for transforming captured content items is provided. The method may include: receiving a capture request for capturing content items in a first application; in response to receiving the capture request, capturing the content items from the first application into a content management tool; applying a generative transformation function to the content items to generate transformed content items; receiving a paste request for pasting the transformed content items into a second application; and in response to receiving the paste request, providing the transformed content items to the second application.

[0007] This summary is provided to present the chosen concepts in a simplified form, which will be further described in the detailed description below. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Additional aspects, features, and / or advantages of the examples will be set forth in part in the description which follows, and in part will be apparent from the description, or may be learned by practice of this disclosure. Attached Figure Description

[0008] Non-limiting and non-exhaustive examples are described with reference to the following figures.

[0009] Figure 1 A block diagram depicts an example of an operating environment in which a content management tool can be implemented, according to the examples of this disclosure; Figure 2A and Figure 2B A flowchart depicts an example method for transforming captured content items according to the examples of this disclosure; Figure 2C A flowchart depicts an example method for transforming captured content items according to the examples of this disclosure; Figures 3A to 3E Screenshots depicting user interface elements of a content management tool according to an example of this disclosure; Figure 4A and Figure 4BAn overview of example generative machine learning models that can be used according to the examples in this disclosure is shown; Figure 5 This is a block diagram illustrating example physical components of a computing device that can implement various aspects of the present disclosure; Figure 6 It is a simplified block diagram of a computing device that can implement various aspects of this disclosure; and Figure 7 This is a simplified block diagram of a distributed computing system in which the various aspects of this disclosure can be implemented. Detailed Implementation

[0010] In the following detailed description, reference is made to the accompanying drawings, which form a part of the description, and specific aspects or examples are illustrated therein by way of illustration. These aspects may be combined, other aspects may be utilized, and structural changes may be made without departing from this disclosure. The aspects may be practiced as methods, systems, or devices. Therefore, the aspects may take the form of hardware implementations, entirely software implementations, or implementations combining software and hardware aspects. Accordingly, the following detailed description should not be considered limiting, and the scope of this disclosure is defined by the appended claims and their equivalents.

[0011] Computing devices encompass a variety of productivity tools and information that facilitate the completion of various tasks, including copying and pasting content items between different devices and applications. For example, clipboard tools allow users to copy and store content items (e.g., images and text) from their original location and paste the copied content items to a new location. However, it can be challenging for users to conveniently and efficiently transform (e.g., translate, correct, adapt, and / or modify) the copied content items before pasting them to a new location.

[0012] According to examples in this disclosure, a content management tool allows a user to capture and generatively transform content items, and copy and paste the transformed content items to a new location. For example, content items may include text, documents, photos, videos, and audio. When a user captures a content item, the content management tool transforms the content item by applying generative transformation features (e.g., translation, correction, adaptation, and / or revision) to transform the content item using a generative large language model (LLM), a transformer model, a diffusion model, or a multimodal model, other types of machine learning models, or a combination of models. For example, a generative transformation feature is a natural language prompt describing one or more tasks to be performed on the content item to generate the transformed content item. The content management tool also presents the transformed content item to the user for further editing and / or copying. In some aspects, a generative transformation feature can be automatically selected based on a previously selected generative transformation feature. Alternatively, a generative transformation feature can be selected from a list of predefined generative transformation features or defined by the user. It should be understood that the captured content item and the transformed content item may be in different modalities.

[0013] According to examples in this disclosure, a content management tool provides user interface elements for user interaction. For example, when a content item is captured, the captured content item is automatically copied into a first user interface element. The content management tool transforms the content item in the first user interface element by applying generative transformation functions defined in a second user interface element, and writes the transformed content item into a third user interface element. The content management tool allows the user to further edit the content in the user interface elements, copy the transformed content item in the third user interface element, and paste the copied content into a new location.

[0014] Figure 1 A block diagram depicts an example of an operating environment 100 in which content management tools can be implemented according to an example of this disclosure. For this purpose, the operating environment 100 includes a computing device 120 associated with a user 110. The operating environment 100 may also include one or more remote devices, such as a productivity platform server 160, communicatively coupled to the computing device 120 via a network 150. The network 150 may include any kind of computing network, including but not limited to wired or wireless local area networks (LANs), wired or wireless wide area networks (WANs), and / or the Internet.

[0015] Computing device 120 includes a content management tool 130 that runs on computing device 120 having a processor 122, memory 124, and communication interface 126. Content management tool 130 allows user 110 to copy, transform, and paste content items. For example, content management tool 130 could be a clipboard or any other productivity tool running on computing device 120 with copy-and-paste and transform capabilities. Content items could be one or more text, documents, images, pictures, photos, videos, or audio. Additionally, computing device 120 could be, but is not limited to, a computer, notebook, laptop, mobile device, smartphone, tablet, portable device, wearable device, or any other suitable computing device capable of running content management tool 130. For this purpose, content management tool 130 also includes a content capture manager 132 and a content transformer 134.

[0016] Content capture manager 132 is configured to receive capture requests for capturing content items. A capture request is any indicator representing a user's intent to capture and generatively transform a content item. A content item can be one or more text, documents, images, pictures, photographs, videos, or audio. Capture requests can be shortcuts and / or gestures assigned by the operating system or by the user. For example, keyboard shortcuts for content capture (e.g., Ctrl+t or Windows logo key+t) can be predefined by the user's operating system and / or by the user. Additionally or alternatively, voice shortcuts for content capture (e.g., "transform the selected content") can be predefined by the user's operating system and / or by the user. Additionally or alternatively, the user can assign gestures as capture requests. For example, the user can specify which screenshot content items they want to capture and transform whenever they take a screenshot on their mobile device. When a capture request is detected, content capture manager 132 is configured to capture the content item and write the captured content item to a user interface element (e.g., an input field) of content management tool 130. In other words, users can define one or more rules or action-based rules as capture requests for capturing content items and copying them into the user interface elements of the content management tool 130.

[0017] An exemplary screenshot of a content management tool 130, including user interface elements for user interaction, is shown in [image / image]. Figure 3A and Figure 3B As shown in the image. Figure 3A As shown, the captured content item is the text string "This is a poorly written text I copid", and in response to being captured, the captured content item is automatically copied into the input field 302 (e.g., the first user interface element) of the content management tool 130.

[0018] Content transformer 134 is configured to apply generative transformation capabilities to captured content items to generate transformed content items using a generative large language model (LLM), a transformer model, a diffusion model or a multimodal model, other types of machine learning models, or a combination of models. As described above, a generative transformation capability is a prompt describing one or more tasks to be performed on a content item to generate a transformed content item. It should be understood that the generative transformation capability to be applied to the content item is presented in a second user interface element (e.g., a prompt field) of the content management tool 130. For example, as... Figure 3A As shown, the generative transformation function to be applied to the captured content item is "Correct English of the INPUT text:", and it is presented in the prompt field 304 of the content management tool 130 (e.g., a second user interface element).

[0019] According to some embodiments, the content transformer 134 is configured to automatically apply a previously selected generative transform function to the captured content item. In some embodiments, the content transformer 134 may receive user input identifying the generative transform function to be applied to the captured content item. For example, a user may select a generative transform function from a predefined list of generative transform functions. Figure 3B As shown, users can select a generative transformation function from a drop-down menu that displays a list of predefined generative transformation functions. Alternatively, users can define generative transformation functions in a prompt field of the content management tool 140. In some embodiments, users can edit existing generative transformation functions presented in the prompt field of the content management tool 140.

[0020] It should be understood that the content prompt database 138 stores one or more predefined generative transformation functions and one or more generative transformation functions that the user has previously used or defined. The content transformer 134 is configured to store previously used generative transformation functions and any edits made and presented to the user. Users can also share one or more generative transformation functions with other users.

[0021] Content transformer 134 is also configured to write the transformed content items into a third user interface element (e.g., an output field) of content management tool 130. For example, as Figure 3A As shown, the original content item "This is a poorly written text I copid" has been corrected to the status "This is a poorly written text I copied", which is presented in the output field 306 of the content management tool 130 (e.g., a third user interface element).

[0022] Once a content item is transformed, the content capture manager 132 is also configured to determine whether an edit request for editing the transformed content item has been received. For example, a user may choose to further edit the transformed content item in the output field. In response to receiving an edit request, the content management tool 130 receives editing of the transformed content item in the output field of the content management tool 130.

[0023] Additionally, the content capture manager 132 is configured to determine whether a copy request has been received to copy the content in the output field of the content management tool 130. It should be understood that the content in the output field of the content management tool 130 is a transformed content item, or, if any editing has been received, an edited transformed content item. In response to receiving a copy request, the content capture manager 132 is configured to store the content in the output field of the content management tool 130 as the final transformed content item in the content database 136. However, it should be understood that in some embodiments, the content management tool 130 may automatically save the transformed content item in the content database 136.

[0024] It should be understood that captured content items are automatically stored in content database 136. It should be understood that content database 136 is synchronized across multiple user devices, allowing the user to capture and paste content items from any of their computing devices. However, it should be understood that, in some respects, content database 136 may be a cloud-based content database shared across multiple user devices.

[0025] Depending on the resources, capabilities, and capacity of the computing device used to capture the content item, the content item can be transformed from the computing device or server 160. For example, if a user captures the content item on their laptop, the content manager tool 130 on the laptop transforms the content item by applying a selected generative transformation function using a generative large language model (LLM), a transformer model, a diffusion model, or a multimodal model, other types of machine learning models, or a combination of models. However, if a user captures the content item on their mobile device, which has fewer resources to perform generative transformations, the content capture manager 132 can send the captured content data to server 160 to transform the content item. The transformed content item is then sent back to the user's mobile device to be inserted into output fields and / or stored in the content database 136.

[0026] Content capture manager 132 is also configured to determine whether a paste request is received to paste copied content at the requested location. The paste request can be a shortcut and / or gesture assigned by the operating system or the user. For example, a keyboard shortcut for content capture (e.g., Ctrl+t or Windows logo key+t) can be assigned by the user's operating system and / or predefined by the user. Additionally or alternatively, a voice shortcut for content capture (e.g., "transform the selected content") can be assigned by the user's operating system and / or predefined by the user. Additionally or alternatively, the user can assign a shortcut or gesture as a paste request. In response to receiving a paste request, content capture manager 132 is configured to paste the most recently captured and transformed content item. It should be understood that the requested location is different from the location from which the content item was originally copied. For example, a user can copy and transform content items from a website and paste the transformed content items into an email. In response to receiving a paste request, content capture manager 132 is configured to write the copied content at the requested location.

[0027] Now for reference Figure 2A and Figure 2B This disclosure provides a method 200 for transforming copied content items according to examples of this disclosure. The general order of the steps in method 200 is as follows: Figure 2A and Figure 2B As shown in the diagram. Typically, method 200 begins at 202 and ends at 232. Method 200 may include more or fewer steps, or may be arranged in a manner consistent with... Figure 2A and Figure 2B The order of the steps shown is different from the order of the steps. For illustrative purposes, method 200 is performed by the computing device of user 110 (e.g., user device 120). However, it should be understood that one or more steps of method 200 may be performed by another device (e.g., server 160).

[0028] Specifically, in some aspects, method 200 can be executed by a content management tool (e.g., 130) running on user device 120. For example, content management tool 130 is a clipboard or other productivity tool running on computing device 120 with copy-and-paste and transform capabilities. For example, computing device 120 can be, but is not limited to, a computer, notebook, laptop, mobile device, smartphone, tablet, portable device, wearable device, or any other suitable computing device capable of executing a content management tool (e.g., 130). For example, server 160 can be any suitable computing device capable of communicating with computing device 120. Method 200 can be executed as a set of computer-executable instructions executed by a computer system and encoded or stored on a computer-readable medium. Furthermore, method 200 can be executed by gates or circuits associated with a processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), system-on-a-chip (SOC), or other hardware device. References will be incorporated herein by reference. Figure 1 and Figure 4 to Figure 7 Method 200 is explained by describing the system, components, modules, software, data structures, user interface, etc.

[0029] Method 200 begins at operation 202, where the process can proceed to 204. At operation 204, content management tool 130 receives a capture request for capturing a content item. A capture request is any indication of a user's intent to capture and generatively transform a content item. The content item can be one or more pieces of text, documents, images, pictures, photos, videos, or audio. The capture request can be a shortcut and / or gesture assigned by the operating system or by the user. For example, a keyboard shortcut for content capture (e.g., Ctrl+t or Windows key+t) can be assigned by the user's computing device's operating system and / or predefined by the user. Additionally or alternatively, a voice shortcut for content capture (e.g., "transform the selected content") can be assigned by the user's computing device's operating system and / or predefined by the user. Additionally or alternatively, the user can assign a gesture as a capture request. For example, the user can specify which screenshot content item the user wants to capture and transform whenever the user takes a screenshot on the user's mobile device.

[0030] At operation 206, in response to receiving a capture request, content management tool 130 captures a content item and automatically provides the captured content item in the input field of content management tool 130. As described above, content management tool 130 provides user interface elements for user interaction. An exemplary screenshot of content management tool 130 including user interface elements for user interaction is shown in... Figure 3A and Figure 3B As shown in the image. Figure 3AAs shown, the captured content item is the text string "This is apoooly written text I copid", and in response to being captured, the captured content item is automatically copied into the input field 302 (e.g., the first user interface element) of the content management tool 130.

[0031] At operation 210, content management tool 130 applies generative transformation functionality to the captured content item to generate a transformed content item using a generative large language model (LLM), a transformer model, a diffusion model or a multimodal model, other types of machine learning models, or a combination of models. As described above, generative transformation functionality is a prompt describing one or more tasks to be performed on the content item to generate the transformed content item. For example, the generative transformation functionality to be applied to the content item is provided in a prompt field (e.g., a second user interface element) of content management tool 130. For example, as... Figure 3A As shown, the generative transformation function to be applied to the captured content item is "Correct English of the INPUT text:", and it is presented in the prompt field 304 of the content management tool 130 (e.g., a second user interface element).

[0032] For this purpose, for example, content management tool 130 can automatically apply the previously selected generative transformation function to the captured content items, as indicated in operation 212.

[0033] In some embodiments, a user can identify generative transformation functions to be applied to captured content items. For example, a user can select a generative transformation function from a list of predefined generative transformation functions, as shown in operation 214. For example, as Figure 3B As shown, a user can select a generative transformation function from a drop-down menu 308 that displays a list of predefined generative transformation functions. In some embodiments, the drop-down menu 308 may also include a predefined number of previously selected generative transformation functions.

[0034] Alternatively, users can define generative transformation functionality in the prompt field of content management tool 140, as shown in operation 216. In some embodiments, users can edit existing generative transformation functionality presented in the prompt field of content management tool 140.

[0035] According to some embodiments, the content management tool 130 may use a machine learning model (e.g., a generative large language model (LLM)) to select or suggest generative transformation features to be applied to captured content items. For example, generative transformation features may be selected or suggested to the user based on one or more generative transformation features previously selected by the user and / or other users for similar types of content items. Additionally, the machine learning model may consider various parameters, including the type of application from which the content item was originally copied, the type of application running on the user's computing device, search history, or any data indicating or suggesting the user's intent for capturing the content item.

[0036] At operation 218, content management tool 130 writes the transformed content item into the output field of content management tool 130. For example, as... Figure 3A As shown, the original content item "This is a poorly written text I copid" has been corrected to the status "This is a poorly written text I copied," which is presented in output field 306 of the content management tool 130 (e.g., a third user interface element). The transformed content item can be one or more texts, documents, images, pictures, photos, videos, or audio. In some embodiments, the modality of the transformed content item differs from that of the original content item. For example, if a user captures the text string "Cat under the Christmas Tree" (i.e., the captured content item), the content management tool 130 can generate an image of a cat under a Christmas tree (i.e., the transformed content item).

[0037] At operation 220, content management tool 130 determines whether an edit request has been received to edit the transformed content item. For example, a user may choose to further edit the transformed content item in the output field. In response to receiving an edit request, content management tool 130 receives the edit to the transformed content item in its output field, as indicated in operation 222.

[0038] At operation 224, content management tool 130 determines whether a copy request has been received to copy the content in the output field of content management tool 130. It should be understood that the content in the output field of content management tool 130 is a transformed content item, or, if any edits have been received at operations 220 to 222, an edited transformed content item.

[0039] At operation 226, in response to receiving a copy request, content management tool 130 stores the content in its output field as the final transformed content item in the database. However, it should be understood that in some embodiments, content management tool 130 may automatically save the transformed content item in the database.

[0040] At operation 228, the content management tool 130 determines whether a paste request has been received to paste the copied content at the requested location. It should be understood that the requested location is different from the location from which the content item was originally copied. For example, a user may copy and transform content items from a website and paste the transformed content items into an email application. In some embodiments, the paste request may be initiated by the operating system or by a user-assigned shortcut and / or gesture. For example, a keyboard shortcut for pasting content (e.g., Ctrl+g or Windows logo key+g) may be initiated by the user's operating system and / or predefined by the user. Additionally or alternatively, a voice shortcut for content capture (e.g., "Paste transformed content") may be initiated by the user's operating system and / or predefined by the user.

[0041] At operation 230, in response to receiving a paste request, content management tool 130 provides the copied content at the requested location. Method 200 can then end at operation 232.

[0042] As described above, the content database is synchronized across multiple devices of the user, allowing the user to capture content items from any of their computing devices. Depending on the resources, capabilities, and capacity of the computing device used to capture the content items, the content items can be transformed from the computing device or server 160. For example, if the user captures content items on their laptop, the content manager tool 130 on the laptop transforms the content items by applying a selected generative transformation function (e.g., translation, correction, adaptation, and / or revision) to use a generative large language model (LLM), a transformer model, a diffusion model, or a multimodal model, other types of machine learning models, or combinations of models. However, if the user captures content items on their mobile device, which has fewer resources to perform generative transformations, the content capture manager 132 can send the captured content data to server 160 to transform the content items. The transformed content items are then sent back to the user's mobile device to be inserted into output fields and / or stored in the content database 136.

[0043] Now for reference Figure 2C This disclosure provides a method 250 for transforming copied content items according to an example. The general order of the steps in method 250 is as follows: Figure 2CAs shown in the diagram. Typically, method 250 begins at 252 and ends at 262. Method 200 may include more or fewer steps, or may be arranged in a manner consistent with... Figure 2C The order of the steps shown differs from the order of the steps in the original text. For illustrative purposes, method 200 is performed by the computing device of user 110 (e.g., user device 120). However, it should be understood that one or more steps of method 200 may be performed by another device (e.g., server 160).

[0044] Specifically, in some aspects, method 250 can be executed by a content management tool (e.g., 130) running on user device 120. For example, content management tool 130 is a clipboard or other productivity tool running on computing device 120 with copy-and-paste and transform capabilities. For example, computing device 120 can be, but is not limited to, a computer, notebook, laptop, mobile device, smartphone, tablet, portable device, wearable device, or any other suitable computing device capable of executing a content management tool (e.g., 130). For example, server 160 can be any suitable computing device capable of communicating with computing device 120. Method 250 can be executed as a set of computer-executable instructions executed by a computer system and encoded or stored on a computer-readable medium. Furthermore, method 200 can be executed by gates or circuits associated with a processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), system-on-a-chip (SOC), or other hardware device. References will be incorporated herein by reference. Figure 1 and Figure 4 to Figure 7 Method 250 is explained by describing the system, components, modules, software, data structures, user interfaces, etc.

[0045] Method 250 begins at operation 252, where the process can proceed to 254. At operation 254, content management tool 130 receives a capture request for capturing a content item. A capture request is any indication of a user's intent to capture and generatively transform a content item. The content item can be one or more text, documents, images, pictures, photos, videos, or audio. The capture request can be a shortcut and / or gesture assigned by the operating system or by the user. For example, a keyboard shortcut for content capture (e.g., Ctrl+t or Windows key +t) can be assigned by the user's computing device's operating system and / or predefined by the user. Additionally or alternatively, a voice shortcut for content capture (e.g., "transform the selected content") can be assigned by the user's computing device's operating system and / or predefined by the user. Additionally or alternatively, the user can assign a gesture as a capture request. For example, the user can specify which screenshot content item the user wants to capture and transform whenever the user takes a screenshot on the user's mobile device.

[0046] At operation 256, in response to receiving a capture request, content management tool 130 captures a content item and applies a generative transformation function to the captured content item to generate a transformed content item using a generative large language model (LLM), a transformer model, a diffusion model or a multimodal model, other types of machine learning models, or a combination of models. As described above, a generative transformation function is a prompt (e.g., a natural language prompt) describing one or more tasks to be performed on the content item to generate the transformed content item.

[0047] In some embodiments, by default, a previously selected generative transformation function (e.g., a generative transformation function used in a previous transformation) is automatically applied to the captured content item to generate a transformed content item. In some embodiments, the content management tool 130 provides user interface elements (e.g., prompt fields) for receiving user input defining the generative transformation functions to be applied to the captured content items. Alternatively, the content management tool 130 provides a drop-down menu with a list of generative transformation functions for the user to select from the drop-down menu. For example, the drop-down menu includes a predefined number of predefined generative transformation functions and / or previously selected generative transformation functions.

[0048] In some embodiments, the content management tool 130 uses a machine learning model (e.g., a generative large language model (LLM)) to select or suggest generative transformation features to be applied to captured content items. For example, generative transformation features may be selected or suggested to the user based on one or more generative transformation features previously selected by the user and / or other users for similar types of content items. Additionally, the machine learning model may consider various parameters, including the type of application from which the content item was originally copied, the type of application running on the user's computing device, search history, or any data indicating or suggesting the user's intent for capturing the content item.

[0049] At operation 258, content management tool 130 determines whether a paste request has been received to paste the transformed content item at the requested location.

[0050] The paste request can be initiated by the operating system or by a user-assigned shortcut and / or gesture. For example, a keyboard shortcut for pasting content (e.g., Ctrl+g or Windows logo key+g) can be initiated by the user's operating system and / or predefined by the user. Additionally or alternatively, a voice shortcut for content capture (e.g., "Paste Transformed Content") can be initiated by the user's operating system and / or predefined by the user. Furthermore, the requested location can differ from the location from which the content item was originally copied. For example, a user can copy and transform a content item from a website and paste the transformed content item into an email application. The transformed content item can be one or more pieces of text, documents, images, pictures, photos, videos, or audio. In some embodiments, the modality of the transformed content item differs from that of the original content item. For example, if a user captures the text string "Cat under the Christmas Tree" (i.e., the captured content item), the content management tool 130 can generate an image of a cat under a Christmas tree (i.e., the transformed content item).

[0051] At operation 260, in response to receiving a paste request, content management tool 130 provides the transformed content item at the requested location. Method 250 can then terminate at operation 262.

[0052] Now for reference Figure 3A and Figure 3B An exemplary screenshot of a content management tool 130, including user interface elements for user interaction, is shown. Figure 3A As shown, the captured content item is the text string "This is a poorly written text I copid", and in response to being captured, the captured content item is automatically copied into the input field 302 of the content management tool 130 (e.g., a first user interface element). Additionally, the generative transformation function to be applied to the captured content item is "Correct English of the INPUT text:", and it is displayed in the prompt field 304 of the content management tool 130 (e.g., a second user interface element).

[0053] In some aspects, generative transformation functions can be automatically selected based on previously chosen generative transformation functions. Alternatively, generative transformation functions can be selected from a list of predefined generative transformation functions or defined by the user. For example, such as Figure 3BAs shown, users can select a generative transformation function from a drop-down menu that displays a list of predefined generative transformation functions. Alternatively, users can define a generative transformation function in the prompt field 304 of the content management tool 140. In some embodiments, users can edit existing generative transformation functions presented in the prompt field 304 of the content management tool 140.

[0054] like Figure 3A As shown, the original content item "This is a poorly written text I copid" has been corrected to the status "This is a poorly written text I copied", which is presented in the output field 306 of the content management tool 130 (e.g., a third user interface element).

[0055] In addition, such as Figure 3B As shown, the content management tool 130 includes a drop-down menu 308 that, when selected, displays a list of predefined generative transformation functions. In some embodiments, the drop-down menu 308 may also include a predefined number of previously selected generative transformation functions.

[0056] Figures 3C to 3E An exemplary screenshot of a content management tool 130 is shown, which includes user interface elements for user interaction, similar to... Figure 3A and Figure 3B The text describes user interface elements with different interface designs. Specifically, Figure 3C and Figure 3D The interface designs 310 and 312 of the content management tool 130 are shown, including a new prompt icon 322, a list 320 with generative transformation functions, and an output field 316 similar to an output field 306. When the user hovers over the new prompt icon 322, a pop-up window 324 appears next to the new prompt icon 322 with the text string "Add new prompt". When the new prompt icon 322 is selected, the interface designs 310 and 312 change to interface design 314, as shown. Figure 3E As shown.

[0057] The generative transformation function list 314 includes a predefined number of predefined generative transformation functions and / or one or more previously selected generative transformation functions. As described above, in the illustrative embodiment, the most recently selected previously selected generative transformation function is automatically selected. Figure 3C and Figure 3DAs shown, "Correct grammar" is automatically selected, and the selected generative transformation function is highlighted. Alternatively, the user can manually select or change the desired generative transformation function from the generative transformation function list 320.

[0058] like Figure 3C As shown, interface design 310 illustrates the situation when there is no transformed content in output field 316. This could be before a capture request is received or after the transformed content has been copied, for example, to the clipboard. It should be understood that in some embodiments, when transformed content (e.g., text) is copied, a pop-up text appears indicating "Text has been copied to the clipboard." When there is no transformed content in output field 316, output field 316 provides a note indicating a shortcut for prompting a capture request and the action to be performed upon receiving the capture request. For example, as... Figure 3C As shown, a comment can state that "when Ctrl+G is pressed, the copied text automatically appears and is transformed according to the selected prompt (e.g., grammar correction)." It should be understood that in some embodiments, the comment can be changed based on the selected generative transformation function and a predefined shortcut used to trigger the copy-and-transform function of the content management tool 130.

[0059] like Figure 3D As shown, interface design 312 illustrates the situation when a capture request is received and the captured content is transformed and provided in output field 316. In this example, the captured content item is the text string "This is a poorly written text I copid", and it is transformed to the current syntax, as selected in the list of generative transformation functions 320. Therefore, the transformed text string "This is a poorly written text I copied" is provided in output field 316. As described above, the user can change the generative transformation function to be applied to the captured content item by selecting one from the list of generative transformation functions 320. Once a generative transformation function is selected, the transformed content item is re-transformed using transformation icon 318. For example, if the user selects the "Translate to Spanish" function and selects transformation icon 318, the content management tool 130 applies the selected generative transformation function to the captured content item and replaces the transformed text string "This is a poorly written text I copied" in output field 318 with the new transformed content.

[0060] Users can select the new notification icon 322 to add a new notification. When the new notification icon 322 is selected, the interface design 314 of the content management tool 130 appears, such as... Figure 3E As shown. The interface design 314 includes input field 328, prompt field 330, and output field 332.

[0061] The hint field 330 indicates the selected generative transformation function. However, the user can define any hints they wish to apply to captured content items. The content management tool 130 allows the user to store user-defined hints (e.g., generative transformation functions) in the content hint database 138 by selecting the save icon 336.

[0062] As described above, the captured content items can be automatically copied to input field 328. Alternatively, the user can manually edit or add content items in input field 328. When icon 334 is selected, content management tool 130 transforms the captured content in input field 328 according to the generative transformation function defined in prompt field 330 to generate and provide the transformed content in output field 332.

[0063] Figure 4A and Figure 4B This provides an overview of example generative machine learning models that can be used based on the aspects described in this article. First, refer to... Figure 4A Conceptual diagram 400 depicts an overview of a pre-trained generative model package 404 according to the aspects described herein, which processes input 402 to generate model output for capturing and generatively transforming content items from generative model output 406 (e.g., transformed content).

[0064] In the example, the generative model package 404 is pre-trained on a variety of inputs (e.g., various human languages, various programming languages, and / or various content types) and therefore does not require fine-tuning or training for a specific scenario. Instead, the generative model package 404 can be pre-trained more generally, such that input 402 includes cues that are generated, selected, or otherwise designed to cause the generative model package 404 to produce certain generative model outputs 406. It should be understood that input 402 and generative model output 406 can each include any of a variety of content types, including but not limited to text output, image output, audio output, video output, programming output, and / or binary output. In the example, input 402 and generative model output 406 can have different content types, as is the case when the generative model package 404 includes a generative multimodal machine learning model.

[0065] Therefore, generative model package 404 can be used in any scenario across a wide range of scenarios, and furthermore, different generative model packages can be used to replace generative model package 404 without substantially modifying other related aspects (e.g., similar to those discussed in this paper). Figure 1 (as described in Figure 3). Therefore, the generative model package 404 operates as a tool for performing machine learning processing, wherein certain inputs 402 of the generative model package 404 are generated programmatically or otherwise determined, thereby enabling the generative model package 404 to produce model outputs 406 that can then be used for further processing.

[0066] Generative model package 404 can be provided or otherwise used according to any of the various paradigms. For example, generative model package 404 can be used on computing devices (e.g., Figure 1 The computing device 140 in the middle can be used locally, or it can be obtained from a machine learning service (e.g., Figure 1 The generative model package 404 is accessed remotely from server 160. In other examples, aspects of the generative model package 404 are distributed across multiple computing devices. In some cases, the generative model package 404 may be accessed via an application programming interface (API), such as by the operating system of the computing device and / or by machine learning services.

[0067] Referring now to aspects of the generative model package 404, the generative model package 404 includes input tokenization 408, input embedding 410, model layer 412, output layer 414, and output decoding 416. In the example, input tokenization 408 processes input 402 to generate input embedding 410, which includes a sequence of symbol representations corresponding to input 402. Therefore, input embedding 410 is processed by model layer 412, output layer 414, and output decoding 416 to produce model output 406. Figure 4B The example architecture corresponding to generative model package 404 is depicted below, which will be discussed in further detail. Even so, it should be understood that the architecture shown and described herein should not be considered limiting, and any of the various other architectures may be used in other examples.

[0068] Figure 4B This is a conceptual diagram depicting an example architecture 450 of a pre-trained generative machine learning model that can be used according to the aspects described herein. As mentioned above, any of the various alternative architectures and corresponding ML models can be used in other examples without departing from the aspects described herein.

[0069] As shown, architecture 450 processes input 402 to produce generative model output 406, the aspects of which are described above regarding... Figure 4AThe discussion took place. Architecture 450 was depicted as a transformer model comprising encoder 452 and decoder 454. Encoder 452 processes input embedding 458 (its aspects can be similar to...). Figure 4A The input embedding (410) includes a sequence of symbolic representations corresponding to input 456. In the example, input 456 includes content data 402 corresponding to a content item.

[0070] Furthermore, positional encoding 460 can incorporate information about the relative and / or absolute positions of the lexical units in the input embedding 458. Similarly, the output embedding 474 includes a sequence of symbolic representations corresponding to the output 472, and positional encoding 476 can similarly incorporate information about the relative and / or absolute positions of the lexical units in the output embedding 474.

[0071] As shown, encoder 452 includes example layer 470. It should be understood that any number of such layers can be used, and the architecture depicted is simplified for illustrative purposes. Example layer 470 includes two sub-layers: a multi-head attention layer 462 and a feedforward layer 466. In the example, residual connections are included around each layer 462, 466, followed by normalization layers 464, 468, respectively.

[0072] Decoder 454 includes example layer 490. Similar to encoder 452, any number of such layers may be used in other instances, and the depicted architecture of decoder 454 is simplified for illustrative purposes. As shown, example layer 490 includes three sublayers: masked multi-head attention layer 478, multi-head attention layer 482, and feedforward layer 486. Aspects of multi-head attention layer 482 and feedforward layer 486 may be similar to those discussed above regarding multi-head attention layer 462 and feedforward layer 466, respectively. Additionally, masked multi-head attention layer 478 performs multi-head attention on the output of encoder 452 (e.g., output 472). In the example, masked multi-head attention layer 478 prevents a position from focusing on subsequent positions. This masking, combined with embedding offsets (e.g., offsetting by one position, as shown in multi-head attention layer 482), ensures that the prediction at a given position depends on the known outputs of one or more positions smaller than the given position. As shown, residual connections are also included around layers 478, 482 and 486, followed by normalized layers 480, 484 and 488, respectively.

[0073] Multi-head attention layers 462, 478, and 482 can each use a set of linear projections for their corresponding dimensions to linearly project the query, key, and value. Attention features (e.g., dot product or additive attention) can be used to process each linear projection, thereby producing [results] for each linear projection. n The resulting values ​​can be concatenated and projected again, so that these values ​​are subsequently displayed as... Figure 4BThe processing is shown (e.g., through the corresponding normalization layer 464, 480 or 484).

[0074] Feedforward layers 466 and 486 can each be fully connected feedforward networks applied to each location. In the example, feedforward layers 466 and 486 each include multiple linear transformations with modified linear unit activations. In the example, each linear transformation is the same across different locations, while different parameters can be used compared to other linear transformations in the feedforward network.

[0075] Furthermore, aspects of the linear transformation 492 can be similar to the linear transformations discussed above regarding the multi-head attention layers 462, 478, and 482 and the feedforward layers 466 and 486. Softmax 494 can also transform the output of the linear transformation 492 into the predicted next lexical probability, as shown in output probability 496. It should be understood that the architectures shown are provided as examples, and in other examples, any of various other model architectures can be used based on the disclosed aspects.

[0076] Therefore, the output probability 496 can be used to form a generative model output 406 according to the aspects described herein, such that the output of the generative ML model (e.g., which may include one or more semantic embeddings and one or more content items) is used as input to determine an action according to the aspects described herein. In other examples, the generative model output 406 is provided as a generated output for transforming captured content items.

[0077] Figures 5 to 7 The associated description provides a discussion of various operating environments in which the aspects of this disclosure can be practiced. However, regarding Figures 5 to 7 The devices and systems shown and discussed are for illustrative purposes and are not intended to limit the wide range of computing device configurations that may be used to practice the aspects of this disclosure described herein.

[0078] Figure 5 This is a block diagram illustrating the physical components (e.g., hardware) of a computing device 500 that can implement various aspects of this disclosure. The computing device components described below can be adapted to the aforementioned computing device, including one or more devices associated with machine learning services (e.g., production platform server 160), and the components mentioned above. Figure 1 The computing device 140 under discussion. In a basic configuration, computing device 500 may include at least one processing unit 502 and system memory 504. Depending on the configuration and type of the computing device, system memory 504 may include, but is not limited to, volatile storage devices (e.g., random access memory), non-volatile storage devices (e.g., read-only memory), flash memory, or any combination of such memories.

[0079] System memory 504 may include an operating system 505 and one or more program modules 506 suitable for running software applications 520, such as one or more components supported by the system described herein. As an example, system memory 504 may store a content capture manager 521 and / or a content converter 522. For instance, operating system 505 may be adapted to control the operation of computing device 500.

[0080] Furthermore, the aspects of this disclosure can be practiced in conjunction with graphics libraries, other operating systems, or any other applications, and are not limited to any particular application or system. This basic configuration is... Figure 5 The components within the dashed line 508 are shown. The computing device 500 may have additional features or functions. For example, the computing device 500 may also include additional data storage devices (removable and / or non-removable), such as disks, optical discs, or magnetic tapes. Such additional storage devices... Figure 5 The image shows a removable storage device 509 and a non-removable storage device 510.

[0081] As described above, multiple program modules and data files can be stored in system memory 504. When executed on processing unit 502, program module 506 (e.g., application 520) can perform processes including, but not limited to, the aspects described herein. Other program modules that can be used according to aspects of this disclosure may include email and contact applications, word processing applications, spreadsheet applications, database applications, PowerPoint presentation applications, drawing or computer-aided applications, etc.

[0082] Furthermore, aspects of this disclosure can be practiced in circuits including discrete electronic components, in packages or integrated electronic chips containing logic gates, in circuits utilizing microprocessors, or on a single chip containing electronic components or a microprocessor. For example, aspects of this disclosure can be practiced via a system-on-a-chip (SOC), wherein... Figure 5 Each or many of the components shown can be integrated onto a single integrated circuit. Such a SoC device may include one or more processing units, graphics units, communication units, system virtualization units, and various application functions, all integrated (or “burned in”) as a single integrated circuit onto a chip substrate. When operating via the SoC, the capabilities described herein regarding the client switching protocol can be operated via dedicated logic integrated with other components of the computing device 500 on the single integrated circuit (chip). Aspects of this disclosure can also be practiced using other techniques capable of performing logical operations (e.g., AND, OR, and NOT), including but not limited to mechanical, optical, fluid, and quantum technologies. Furthermore, aspects of this disclosure can be practiced within a general-purpose computer or in any other circuit or system.

[0083] The computing device 500 may also have one or more input devices 512, such as a keyboard, mouse, pen, voice or speech input device, touch or swipe input device, etc. It may also include output devices 514, such as a display, speaker, printer, etc. The above devices are examples, and other devices may be used. The computing device 500 may include one or more communication connections 516 that allow communication with other computing devices 550. Examples of suitable communication connections 516 include, but are not limited to, radio frequency (RF) transmitters, receivers, and / or transceiver circuitry; universal serial buses (USB), parallel and / or serial ports.

[0084] As used herein, the term computer-readable medium can include computer storage media. Computer storage media can include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, or program modules. System memory 504, removable storage device 509, and non-removable storage device 510 are examples of computer storage media (e.g., memory storage devices). Computer storage media can include RAM, ROM, electrically erasable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical storage devices, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other article of manufacture that can be used to store information and can be accessed by computing device 500. Any such computer storage medium may be part of computing device 500. Computer storage media does not include carrier waves or other propagated or modulated data signals.

[0085] Communication media can be embodied in computer-readable instructions, data structures, program modules, or other data (such as carrier waves or other transmission mechanisms) in modulated data signals, and include any information transmission medium. The term "modulated data signal" can describe a signal having one or more characteristics set or altered in a manner that encodes information in the signal. By way of example and not limitation, communication media can include wired media such as wired networks or direct wired connections, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.

[0086] Figure 6System 600 is illustrated, which can be, for example, a mobile computing device such as a mobile phone, smartphone, wearable computer (such as a smartwatch), tablet computer, laptop computer, etc., and various aspects of this disclosure can be implemented using this system. In one example, system 600 is implemented as a "smartphone" capable of running one or more applications (e.g., browser, email, calendar, contact manager, messaging client, game, and media client / player). In some aspects, system 600 is integrated as a computing device, such as an integrated personal digital assistant (PDA) and cordless phone.

[0087] In its basic configuration, this type of mobile computing device is a handheld computer with both input and output elements. System 600 typically includes a display 605 and one or more input buttons that allow the user to input information into system 600. The display 605 can also be used as an input device (e.g., a touchscreen display).

[0088] If included, optional side input elements allow for further user input. For example, a side input element could be a rotary switch, a button, or any other type of manual input element. Alternatively, system 600 may include more or fewer input elements. For example, in some aspects, display 605 may not be a touchscreen. In another example, an optional keypad 635 may also be included, which could be a physical keypad or a “soft” keypad generated on a touchscreen display.

[0089] In various aspects, output elements include a display 605 for displaying a graphical user interface (GUI), a visual indicator (e.g., a light-emitting diode 620), and / or an audio transducer 625 (e.g., a speaker). In some aspects, a vibration transducer is included to provide tactile feedback to the user. In yet another aspect, input and / or output ports are included, such as audio inputs (e.g., a microphone jack), audio outputs (e.g., a headphone jack), and video outputs (e.g., an HDMI port), for sending signals to or receiving signals from external devices.

[0090] One or more applications 666 may be loaded into memory 662 and run on or associated with operating system 664. Examples of applications include telephone dialers, email programs, personal information management (PIM) programs, word processing programs, spreadsheet programs, internet browser programs, messaging programs, etc. System 600 also includes a non-volatile storage area 668 within memory 662. The non-volatile storage area 668 may be used to store persistent information that should not be lost if system 600 is powered off. Applications 666 may use and store information in the non-volatile storage area 668, such as emails or other messages used by email applications. A synchronization application (not shown) also resides on system 600 and is programmed to interact with a corresponding synchronization application residing on a host computer to keep the information stored in the non-volatile storage area 668 synchronized with the corresponding information stored on the host computer. It should be understood that other applications may be loaded into memory 662 and run on system 600 as described herein (e.g., content capture managers, content converters, etc.).

[0091] System 600 has a power supply 670, which can be implemented as one or more batteries. The power supply 670 may also include an external power source, such as an AC adapter or a power docking station for replenishing or recharging the batteries.

[0092] System 600 may also include a radio interface layer 672 that performs functions for transmitting and receiving radio frequency communications. Radio interface layer 672 facilitates wireless connectivity between system 600 and the "external world" via a communications operator or service provider. Transmissions to and from radio interface layer 672 are conducted under the control of operating system 664. In other words, communications received by radio interface layer 672 can be propagated to application 666 via operating system 664, and vice versa.

[0093] A visual indicator 620 can be used to provide visual notifications, and / or an audio interface 674 can be used to generate audible notifications via an audio transducer 625. In the illustrated example, the visual indicator 620 is a light-emitting diode (LED), and the audio transducer 625 is a speaker. These devices can be directly coupled to a power supply 670 such that when activated, they remain on for a duration indicated by the notification mechanism, even if the processor 660 and other components may be turned off to conserve battery power. The LED can be programmed to remain on indefinitely until the user takes action to indicate the device's power-on status. The audio interface 674 is used to provide and receive audible signals to and from the user. For example, in addition to being coupled to the audio transducer 625, the audio interface 674 can also be coupled to a microphone to receive audible input, such as to facilitate telephone conversations. According to various aspects of this disclosure, the microphone can also be used as an audio sensor to facilitate control of notifications, as described below. The system 600 may also include a video interface 676, which enables the operation of the onboard camera 630 to record still images, video streams, etc.

[0094] It should be understood that system 600 may have additional features or functions. For example, system 600 may also include additional data storage devices (removable and / or non-removable), such as disks, optical discs, or magnetic tapes. Such additional storage devices... Figure 6 The non-volatile storage region 668 is shown in the middle.

[0095] As described above, data / information generated, captured, and stored via system 600 can be stored locally, or the data can be stored on any number of storage media, which can be accessed by the device via radio interface layer 672 or via a wired connection between system 600 and a separate computing device associated with system 600 (e.g., a server computer in a distributed computing network such as the Internet). It should be understood that such data / information can be accessed via radio interface layer 672 or via a distributed computing network. Similarly, depending on any of the various data / information transmission and storage components, including email and collaborative data / information sharing systems, such data / information can be easily transferred between computing devices for storage and use.

[0096] Figure 7 One aspect of the architecture of a system for processing data received at a computing system from a remote source, such as a personal computer 704, a tablet computing device 706, or a mobile computing device 708, is illustrated, as described above. The content displayed at server device 702 can be stored in different communication channels or other storage types. For example, various documents can be stored using a directory service 724, a web portal 725, an email service 726, an instant messaging repository 728, or a social networking site 730.

[0097] Application 720 (e.g., similar to application 520) can be adopted by a client communicating with server device 702. Additionally or alternatively, server device 702 may employ content capture manager 791 and / or content transformer 792. Server device 702 can provide data to and from client computing devices such as personal computer 704, tablet computing device 706, and / or mobile computing device 708 (e.g., smartphone) via network 715. As an example, the aforementioned computer system may be embodied in personal computer 704, tablet computing device 706, and / or mobile computing device 708 (e.g., smartphone). In addition to receiving graphics data that can be preprocessed at the graphics initiating system or post-processed at the receiving computing system, any of these examples of computing devices may also obtain content from repository 716.

[0098] It should be understood that the aspects and functions described herein can operate on distributed systems (e.g., cloud-based computing systems), where application functions, memory, data storage and retrieval, and various processing functions can remotely operate on each other via distributed computing networks (such as the Internet or intranets). Various types of user interfaces and information can be displayed via onboard computing device displays or via remote display units associated with one or more computing devices. For example, various types of user interfaces and information can be displayed and interacted with on a wall surface on which various types of user interfaces and information are projected. Interaction with multiple computing systems on which aspects of this disclosure can be practiced includes keystroke input, touchscreen input, voice or other audio input, gesture input, wherein the associated computing device is equipped with detection (e.g., camera) functions for capturing and interpreting user gestures to control the functions of the computing device, etc.

[0099] For example, the aspects of this disclosure have been described above with reference to block diagrams and / or operational illustrations of methods, systems, and computer program products according to various aspects of this disclosure. The functions / actions marked in the boxes may not occur in the order shown in any flowchart. For example, two boxes shown consecutively may actually be executed substantially simultaneously, or these boxes may sometimes be executed in reverse order, depending on the functions / actions involved.

[0100] The description and illustrations of one or more aspects provided in this application are not intended to limit or restrict the scope of the claimed disclosure in any way. The aspects, examples, and details provided in this application are considered sufficient to convey ownership and enable others to make and use the claimed aspects of this disclosure. The claimed disclosure should not be construed as limited to any aspect, example, or detail provided in this application. Whether shown and described in combination or separately, various features (both structural and methodological) are intended to be selectively included or omitted to produce aspects having a particular set of features. Given the descriptions and illustrations provided in this application, those skilled in the art can contemplate variations, modifications, and alternatives falling within the spirit of the broader aspects of the overall inventive concept embodied in this application without departing from the broader scope of the claimed disclosure.

[0101] Furthermore, the aspects and functions described herein can operate on distributed systems (e.g., cloud-based computing systems), where application functions, memory, data storage and retrieval, and various processing functions can remotely operate on each other via distributed computing networks (such as the Internet or intranets). Various types of user interfaces and information can be displayed via onboard computing device displays or via remote display units associated with one or more computing devices. For example, various types of user interfaces and information can be displayed and interacted with on a wall surface on which various types of user interfaces and information are projected. Interaction with multiple computing systems on which aspects of this disclosure can be practiced includes keystroke input, touchscreen input, voice or other audio input, gesture input, wherein the associated computing device is equipped with detection (e.g., camera) functions for capturing and interpreting user gestures to control the functions of the computing device, etc.

[0102] The phrases “at least one,” “one or more,” “or,” and “and / or” are open-ended expressions that are both connected and separate in operation. For example, each of the expressions “at least one of A, B, and C,” “at least one of A, B, or C,” “one or more of A, B, and C,” “one or more of A, B, or C,” “A, B, and / or C,” and “A, B, or C” means A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together.

[0103] The term "a" or "an" entity refers to one or more of the same entity. Therefore, the terms "a," "one or more," and "at least one" are used interchangeably herein. It should also be noted that the terms "comprising," "including," and "having" are used interchangeably.

[0104] As used herein, the term "automatic" and its variations refer to any process or operation that is typically continuous or semi-continuous and can be performed without substantial human input. However, a process or operation can be automatic if input is received prior to its execution, even if the execution of the process or operation uses substantial or non-substantial human input. Human input is considered substantial if it influences how the process or operation will be performed. Human input agreeing to perform a process or operation is not considered "substantial."

[0105] Any steps, functions, and operations discussed in this article can be performed continuously and automatically.

[0106] Example systems and methods of this disclosure have been described with respect to computing devices. However, to avoid unnecessarily obscuring this disclosure, several known structures and devices have been omitted from the foregoing description. Such omissions should not be construed as limiting. Specific details have been set forth to provide an understanding of this disclosure. However, it should be understood that this disclosure can be practiced in various ways beyond the specific details set forth herein.

[0107] Furthermore, while the examples shown herein illustrate various components of a co-located system, some components may be located remotely in distant parts of a distributed network (such as a LAN and / or the Internet), or within a dedicated system. Therefore, it should be understood that system components may be combined into one or more devices, such as servers, communication equipment, or co-located on specific nodes of a distributed network, such as analog and / or digital telecommunications networks, packet-switched networks, or circuit-switched networks. As can be understood from the foregoing description, and for computational efficiency reasons, system components can be positioned anywhere within the distributed network of components without affecting the operation of the system.

[0108] Furthermore, it should be understood that the various links connecting the elements can be wired or wireless links, or any combination thereof, or any other known or later-developed element capable of providing data to and / or transmitting data from the connected elements. These wired or wireless links can also be secure links and capable of transmitting encrypted information. For example, the transmission medium used as a link can be any suitable carrier for electrical signals, including coaxial cables, copper wires, and optical fibers, and can take the form of sound waves or light waves, such as those generated during radio wave and infrared data communication.

[0109] Although flowcharts have been discussed and explained regarding specific event sequences, it should be understood that changes, additions, and omissions to the sequence can occur without substantially affecting the operation of the disclosed configurations and aspects.

[0110] Several variations and modifications of this disclosure may be used. Some features of this disclosure may be provided without providing others.

[0111] In another configuration, the systems and methods of this disclosure may be implemented in combination with a dedicated computer, a programmable microprocessor or microcontroller and peripheral integrated circuit elements, an ASIC or other integrated circuit, a digital signal processor, hardwired electronic or logic circuitry (such as discrete component circuitry), a programmable logic device or gate array (such as a PLD, PLA, FPGA, PAL), a dedicated computer, any similar components, etc. Generally, any device or apparatus capable of implementing the methods shown herein can be used to implement various aspects of this disclosure. Example hardware that may be used in this disclosure includes computers, handheld devices, telephones (e.g., cellular, internet-enabled, digital, analog, hybrid, etc.), and other hardware known in the art. Some of these devices include processors (e.g., single or multiple microprocessors), memory, non-volatile storage devices, input devices, and output devices. Furthermore, alternative software implementations, including but not limited to distributed processing or component / object distributed processing, parallel processing, or virtual machine processing, may also be constructed to implement the methods described herein.

[0112] In another configuration, the disclosed method can be readily implemented using software from an object-oriented or object-based software development environment that provides portable source code usable on various computer or workstation platforms. Alternatively, the disclosed system can be implemented partially or entirely in hardware using standard logic circuitry or VLSI design. Whether to implement the system according to this disclosure using software or hardware depends on the system's speed and / or efficiency requirements, specific functions, and the particular software or hardware system or microprocessor or microcomputer system used.

[0113] In another configuration, the disclosed method can be implemented in part in software, which can be stored on a storage medium and executed on a programmed general-purpose computer in cooperation with a controller and memory, a dedicated computer, a microprocessor, etc. In these cases, the systems and methods of this disclosure can be implemented as programs embedded in a personal computer, such as applets in JAVA® or CGI scripts, as resources residing on a server or computer workstation, or as routines embedded in a dedicated measurement system, system component, etc. The system can also be implemented by physically integrating the system and / or method into a software and / or hardware system.

[0114] If described herein, this disclosure is not limited to standards and protocols. Other similar standards and protocols not mentioned herein exist and are included in this disclosure. Furthermore, the standards and protocols mentioned herein, as well as other similar standards and protocols not mentioned herein, are periodically replaced by faster or more efficient equivalents with substantially the same functionality. Such alternative standards and protocols with the same functionality are considered equivalents included in this disclosure.

[0115] According to at least one example of this disclosure, a method for transforming captured content items is provided. The method may include: receiving a capture request for capturing the content item; upon receiving the capture request, capturing the content item and providing the content item in a first user interface element of a content management tool; applying a generative transformation function to the content item to generate a transformed content item; writing the transformed content item into a second user interface element of the content management tool; receiving a paste request for pasting the transformed content item at a requested location; and in response to receiving the paste request, providing the transformed content item at the requested location.

[0116] According to at least one aspect of the above method, the method may include: receiving a capture request for capturing a content item includes receiving a capture request for capturing a content item in a first application, and wherein the requested location is in a second application different from the first application.

[0117] According to at least one aspect of the above method, the method may include applying a generative transformation function to a content item by: automatically applying a previously selected generative transformation function to the content item.

[0118] According to at least one aspect of the above method, the method may include applying a generative transformation function to a content item by: receiving user input indicating that the generative transformation function should be applied to the content item.

[0119] According to at least one aspect of the above method, the method may include an indication of a generative transformation function in a third user interface element of a content management tool.

[0120] According to at least one aspect of the above method, the method may include a selection of a generative transformation function from a list of predefined generative transformation functions by user input.

[0121] According to at least one aspect of the above method, the method may include: wherein the generative transformation function is a natural language prompt that describes one or more tasks to be performed on the content item to generate the transformed content item.

[0122] According to at least one aspect of the above method, the method may further include: receiving an edit request for editing the transformed content item before receiving a copy request.

[0123] According to at least one aspect of the above method, the method may further include: receiving a copy request for copying the transformed content item; and in response to receiving the copy request, storing the transformed content item in a database.

[0124] According to at least one aspect of the above method, the method may include: applying a generative transformation function to a content item to generate a transformed content item includes: applying the generative transformation function to the content item using at least one of the following: a generative large language model (LLM), a transformer model, a diffusion model, or a multimodal model.

[0125] According to at least one aspect of the above method, the method may include: wherein the content item is at least one of text, image or audio, and the transformed content item is at least one of text, image or audio.

[0126] According to at least one example of this disclosure, a computing device for transforming captured content items is provided. The computing device may include: a processor; and a memory storing a plurality of instructions, which, when executed by the processor, cause the computing device to: receive a capture request for capturing content items; in response to the capture request, capture the content items and provide the content items in a first user interface element of a content management tool; apply a generative transformation function to the content items to generate transformed content items; write the transformed content items into a second user interface element of the content management tool; receive a paste request for pasting the transformed content items at a requested location; and in response to the paste request, provide the transformed content items at the requested location.

[0127] According to at least one aspect of the computing device described above, the computing device may include: receiving a capture request for capturing content items includes: receiving a capture request for capturing content items in a first application, and wherein the requested location is in a second application different from the first application.

[0128] According to at least one aspect of the computing device described above, the computing device may include applying generative transformation functions to content items by automatically applying previously selected generative transformation functions to the content items.

[0129] According to at least one aspect of the computing device described above, the computing device may include: wherein applying the generative transformation function to a content item includes: receiving user input indicating that the generative transformation function should be applied to the content item.

[0130] According to at least one aspect of the aforementioned computing device, the computing device may include an instruction for a generative transformation function in a third user interface element of a content management tool, or a selection of a generative transformation function from a list of predefined generative transformation functions.

[0131] According to at least one aspect of the computing device described above, the computing device may include: wherein the generative transformation function is a natural language prompt that describes one or more tasks to be performed on the content item to generate the transformed content item.

[0132] According to at least one example of this disclosure, a method for transforming captured content items is provided. The method may include: receiving a capture request for capturing content items in a first application; in response to receiving the capture request, capturing the content items from the first application into a content management tool; applying a generative transformation function to the content items to generate transformed content items; receiving a paste request for pasting the transformed content items into a second application; and in response to receiving the paste request, providing the transformed content items to the second application.

[0133] According to at least one aspect of the above method, the method may include: wherein the generative transformation function is a natural language prompt that describes one or more tasks to be performed on the content item to generate the transformed content item.

[0134] According to at least one aspect of the above method, the method may include applying a generative transformation function to a content item by: automatically applying a previously selected generative transformation function to the content item; or receiving user input indicating that a generative transformation function should be applied to the content item.

[0135] In various configurations and aspects, this disclosure includes components, methods, processes, systems, and / or apparatuses substantially as depicted and described herein, including various combinations, sub-combinations, and subsets thereof. Upon understanding this disclosure, those skilled in the art will understand how to make and use the systems and methods disclosed herein. In various configurations and aspects, this disclosure includes providing apparatus and processes in the absence of items not depicted and / or described herein, including in the absence of such items that may have already been used in prior apparatus or processes, for example, to improve performance, simplify implementation, and / or reduce implementation costs.

Claims

1. A method for transforming captured content items, the method comprising: Receive capture requests used to capture content items; Upon receiving the capture request, the content item is captured and provided in the first user interface element of the content management tool; Apply the generative transformation function to the content item to generate the transformed content item; The transformed content items are written into the second user interface element of the content management tool; Receive a paste request to paste the transformed content item at the requested location; as well as In response to receiving the paste request, the transformed content item is provided at the requested location.

2. The method according to claim 1, Receiving the capture request for capturing the content item includes receiving the capture request for capturing the content item in the first application, and The requested location is in a second application, which is different from the first application.

3. The method according to claim 1, wherein applying the generative transformation function to the content item includes: The previously selected generative transformation function is automatically applied to the content item.

4. The method of claim 1, wherein applying the generative transformation function to the content item comprises: Receive user input, which indicates that the generative transformation function should be applied to the content item.

5. The method of claim 4, wherein the user input is an indication of the generative transformation function in a third user interface element of the content management tool.

6. The method of claim 4, wherein the user input is a selection of the generative transformation function from a list of predefined generative transformation functions.

7. The method of claim 1, wherein the generative transformation function is a natural language prompt that describes one or more tasks to be performed on the content item to generate the transformed content item.

8. The method according to claim 1, further comprising: Before receiving a copy request, an edit request is received for editing the transformed content item.

9. The method according to claim 1, further comprising: Receive a copy request for copying the transformed content item; as well as In response to receiving the copy request, the transformed content item is stored in the database.

10. The method of claim 1, wherein applying the generative transformation function to the content item to generate the transformed content item comprises: Use at least one of the following to apply the generative transformation function to the content item: generative large language model (LLM), transformer model, diffusion model, or multimodal model.

11. The method of claim 1, wherein the content item is at least one of text, image, or audio, and the transformed content item is at least one of text, image, or audio.

12. A computing device for transforming captured content items, the computing device comprising: Processor (122); as well as A memory (124) storing a plurality of instructions, which, when executed by the processor, cause the computing device (122) to: Receive capture requests used to capture content items; In response to the capture request, the content item is captured and provided in a first user interface element of the content management tool; Apply the generative transformation function to the content item to generate the transformed content item; The transformed content items are written into the second user interface element of the content management tool; Receive a paste request to paste the transformed content item at the requested location; as well as In response to the paste request, the transformed content item is provided at the requested location.

13. The computing device of claim 12, wherein receiving the capture request for capturing the content item comprises: Receive a capture request for capturing content items in a first application, wherein the requested location is in a second application different from the first application.

14. The computing device of claim 12, wherein applying the generative transformation function to the content item comprises: The previously selected generative transformation function is automatically applied to the content item.

15. The computing device of claim 12, wherein applying the generative transformation function to the content item comprises: Receive user input, which indicates that the generative transformation function should be applied to the content item.

16. The computing device of claim 15, wherein the user input is an indication of the generative transformation function in a third user interface element of the content management tool, or the user input is a selection of the generative transformation function from a list of predefined generative transformation functions.

17. The computing device of claim 12, wherein the generative transformation function is a natural language prompt describing one or more tasks to be performed on the content item to generate the transformed content item.

18. A method for transforming captured content items, the method comprising: Receive capture requests for capturing content items in the first application; In response to receiving the capture request, the content item is captured from the first application and transferred to the content management tool; Apply the generative transformation function to the content item to generate the transformed content item; Receive a paste request for pasting the transformed content item into a second application; as well as In response to receiving the paste request, the transformed content item is provided to the second application.

19. The method of claim 18, wherein the generative transformation function is a natural language prompt that describes one or more tasks to be performed on the content item to generate the transformed content item.

20. The method of claim 18, wherein applying the generative transformation function to the content item comprises: The previously selected generative transformation function is automatically applied to the content item; or Receive user input, which indicates that the generative transformation function should be applied to the content item.